The documentary case file
Part of the AI-Native Shift

7%

of leaders report established AI ROI.

KPMG Global AI Pulse, Q2 2026

Adoption is nearly universal. Proven return is almost nonexistent. The gap between those two facts is the whole story — and it is not a measurement problem you can buy your way out of.

The fog was never the number. It was the visibility you never had.

Cold open

One number opens the case.

Boards are spending on AI and cannot find the return. Everyone is asking what the return on AI is as if it were a new question. It is the oldest question in the business — and the industry is about to make the exact same mistake with AI that it made with marketing attribution.

Before the case can build to an answer, it has to take apart the numbers most leaders reach for first — the ones that feel like proof and measure nothing that matters.

The autopsy

Five metrics that lie, examined and struck.

Most AI measurement tracks whether people are using the tool. That is like measuring whether employees showed up to the office and calling it productivity. Here is each metric a dashboard is proudest of, struck through, with the metric worth watching in its place — the reframe is the measurement engine’s own.

Struck from the case
Prompts per user per day

Decisions improved per engagement — did the AI-assisted work change the outcome?

Measures activity, not outcomes.

Struck from the case
Time saved per task

Tasks eliminated or transformed — work that no longer exists because AI changed the approach entirely.

Saving 4 hours on a report nobody reads is faster waste.

Struck from the case
Adoption rate / active users

Output quality delta — is the work product measurably better than before?

Tells you people are logging in. It says nothing about whether the work coming out the other side is meaningfully different.

Struck from the case
Token usage / API calls

Cost per outcome — what did each meaningful business result cost in AI resources?

Infrastructure cost data disguised as value data.

“I’ve seen companies spending $50,000 a month on tokens. That’s not the flex you think it is.”

Erin Wiggers, on Value-First Platform: AI Data Readiness

Struck from the case
Emails sent / content generated

Response quality and engagement depth — did the AI-assisted communication produce a different kind of response?

Volume is the Measurement Trap in its purest form. More is not better. More is just more.

The turn

You don’t need to do attribution, you just need visibility. These are very different things.”

Chris Carolan, on Value-First Platform: AI Data Readiness

That is the reframe that clears the fog. You were never going to attribute a dollar to a prompt. But visibility into how the work actually flows — which engagement, what the AI traceably did, whether a detection caused the save — is a different thing entirely, and it is the thing that gives you an honest shot at a real number. Look at the board: the answer was there the whole time. Three dimensions, not one.

Dimension 1
Revenue Influenced

The revenue of deals where AI agents were traceably part of the engagement, tied to specific deal activity through a shared system.

The guard: Not every deal where someone used ChatGPT — the AI must be traceable to the deal, not self-reported.

Dimension 2
Revenue Protected

Revenue retained or expanded where an AI-detected signal caused the human action — a health score, a sentiment shift, a risk flag someone acted on.

The guard: Only when the AI detection triggered the action. If the team would have caught it anyway, it does not count.

Dimension 3
Capability Built

What the team can do now that it could not six months ago — decisions it could not have made, work it could not have produced. The hardest to measure and the most valuable.

The guard: Not hours saved. A team spending one hour to produce requirements it could not produce before did not save nineteen hours; it made an informed decision instead of showing up unprepared.

The instrument

The number carries only the confidence you have earned.

Before you build anything, place yourself. These five levels are not a score — they set how much weight the number is allowed to bear. Where you honestly sit decides whether the answer is trustworthy at all, directional, or answerable with a real figure.

No Visibility
Not yet trustworthy
AI tools are in use, but nothing is tracked. Asked for the ROI, the answer is anecdotal or silence.
Activity Tracking
Not yet trustworthy
You know who uses AI and how often — but cannot connect any of it to a business outcome.
Outcome Awareness
Directional
You can point to specific outcomes AI influenced, but the evidence is curated, not systematic.
Systematic Measurement
Answerable
All three dimensions are tracked through integrated systems. You can answer the question with a real number.
Compound Intelligence
Answerable
The measurement system itself uses AI to find where to deploy AI next.

The level is not a score. It is the confidence the number below is allowed to carry: not yet trustworthy — you have activity, not outcome visibility. at the first two levels, directional — curated, not comprehensive. at the third, and answerable with a number that connects to real revenue and real capability. at the top two.

The confidence tiers those levels map to

Not yet trustworthyLevels 0 to 1
not yet trustworthy — you have activity, not outcome visibility.
DirectionalLevel 2
directional — curated, not comprehensive.
AnswerableLevels 3 to 4
answerable with a number that connects to real revenue and real capability.

This is not a meter to fill. Each tier is the confidence a number is allowed to carry at that level of visibility — nothing more. A directional number you can defend is worth more than an answerable one you cannot, and reaching a higher tier is never the point. Honest measurement is. Chasing the top tier for its own sake is the same trap as chasing the return: it makes the label the goal instead of the value underneath it.

Build your case file

Now assemble it with your own numbers.

Enter your AI spend, the three dimensions, and your readiness level. The measurement is arithmetic over your own entries and nothing else. If your visibility has not earned a number yet, this instrument tells you so and hands you the baseline to build instead — it will not invent one.

Your workbench
The case is made. Now build yours — the measurement is arithmetic over your entries and nothing else. It assembles in your case file as you enter it.Watch it assemble
Context

AI spend for the period

The denominator — measured against value created, never a value metric on its own.

Dimension 1

Revenue Influenced

Deals AI was traceably part of — tied to the deal, not “someone used ChatGPT once.” Describe what the AI traceably did in your own words, and be honest: this is the question your board will ask, and the tool records your attestation rather than judging it. An entry with the engagement or that description left blank is shown, not counted.

Entry 1
Dimension 2

Revenue Protected

Retention or expansion an AI-detected signal caused. If the team would have caught it anyway, it does not count.

Entry 1
Did the AI detection cause the human action?
Dimension 3

Capability Built

What the team can do now that it could not six months ago. Not a dollar figure. Not hours saved.

Entry 1
Confidence

Your readiness level

Be honest here — it sets the confidence the number is allowed to carry, and the tool checks your claim against the evidence you entered.

The level is not a score. It is the confidence the number below is allowed to carry: not yet trustworthy — you have activity, not outcome visibility. at the first two levels, directional — curated, not comprehensive. at the third, and answerable with a number that connects to real revenue and real capability. at the top two.

The refusal

The metrics that do not count

Try to add hours saved, adoption, tokens, or any other metric. The tool refuses each with the metric worth watching instead.

Nothing here reaches the value read. The value picture is built only from Revenue Influenced, Revenue Protected, and Capability Built.

Select a readiness level to build the measurement.

Both panels in full, in the page. They open over the argument as fly-ins when you summon them; this is the same content, always here, whether or not that layer is available to you.

The Case Board

7%

of leaders report established AI ROI.

KPMG Global AI Pulse, Q2 2026

The case builds to two revenue figures and a capability — never one blended number.

Where you are

Cold Open

Read the reckoning, then watch five metrics get taken apart.

Your case file

Empty — nothing entered yet

The denominator

AI spend for the period — the denominator, never a value metric on its own.

What the case builds to

Revenue Influenced

Traceable AI-assisted revenue — tied to the deal, not self-reported.

Revenue Protected

Retention or expansion an AI-detected signal caused — not a save you would have made anyway.

Capability Built

What the team can do now that it could not before — ordinal, never dollarized.

Two revenue figures and a capability — never one blended number.

Excluded from your file

The five metrics that lie never enter the case — the engine measures none of them:

  • Prompts per user per day
  • Time saved per task
  • Adoption rate / active users
  • Token usage / API calls
  • Emails sent / content volume

The confidence your number can carry

Where you honestly sit sets how much weight the number can bear.

No Visibility
Not yet trustworthy
Activity Tracking
Not yet trustworthy
Outcome Awareness
Directional
Systematic Measurement
Answerable
Compound Intelligence
Answerable

The level is not a score. It is the confidence the number below is allowed to carry: not yet trustworthy — you have activity, not outcome visibility. at the first two levels, directional — curated, not comprehensive. at the third, and answerable with a number that connects to real revenue and real capability. at the top two.

Not yet trustworthyLevels 0 to 1
DirectionalLevel 2
AnswerableLevels 3 to 4

Enter your own AI spend, the Three Dimensions, and your readiness level, then build the measurement. The picture that follows is arithmetic over your entries — nothing else.

Keep going

The Clearing builds the case. Two more doors open onto the same instrument, and the full framework goes deeper on every part.

Start here
Ground Truth

Set the baseline first. See whether your number can bear weight, or whether the honest next step is a baseline, not a return.

Enter Ground Truth
Confront it
The Reckoning Room

Present your case to an adversarial board and watch it deflate, entry by entry, to the number that survives.

Enter the Reckoning Room
Read it
The full framework

The deep-dive: the five metrics that lie, the three dimensions in full, all five readiness levels, and the four prerequisites.

Open the deep-dive
See where AI ROI sits on the Value Path →

Framework developed in collaboration with Trisha Merriam and Erin Wiggers on the Value-First Platform: AI Data Readiness show.