June 10, 2026 · 7 min read · Measuring the payoff
Why your AI ROI number is wrong
Not because AI does not work. Because the usual way of building the number holds at least eight known errors, and each one pushes the number up.
Your AI return number is probably too high. The tools do help, and the counted gains in work finished are real — if smaller than the ads say. The problem is how the number gets built. The standard sum stacks up known errors that all push the same way, and none of them cancel out.
Here they are, named, with what the proof says about each. Then the useful half: what to report instead.
KPMG's Global AI Pulse for Q2 2026 asked 2,145 C-suite leaders in 20 countries. Only 7% say they have proven a return on AI: 15% among leaders who can see their AI costs clearly, 3% among those who cannot. That gap is not about sharper math. It is about whether the inputs could be trusted.
Eight ways the number breaks
1. Counting the same saved hour twice
An engineer saves an hour. That hour gets booked three times: as time saved in the work model, as faster features in the roadmap case, and as a hire you never had to make. Three lines, three dollar amounts, one hour.
The fix is to count each dollar once, in a bucket you pick up front. Claim the hour as a hire you skipped and you cannot also claim it as new features. Most inflated models break here before they break anywhere else.
2. No number for what the speed cost you
Each speed number needs a second number beside it, and that second number is the cost of going faster. Report how fast work merged but not how much of it had to be redone, and you have half a picture. The flattering half.
Faros AI's analysis of 10,000+ developers across 1,255 teams (July 2025) found AI users finished 21% more tasks and merged 98% more pull requests. It also found 9% more bugs per developer, 154% larger pull requests, and 91% more time spent reviewing them. It found no clear link between AI use and gains at the company level. Opsera's 2026 AI Coding Impact Benchmark, drawing on 250,000+ developers, says AI-written pull requests take 4.6x longer to review. That one is published by a vendor and its method is not fully public, so weigh it with that in mind. It does point the same way as the independent Faros data.
Work that moves from the writer to the reviewer has not gone away. It moved to a different budget.
3. The teams chose the tool themselves
You compared heavy AI teams against light ones and found a gap. But the teams that took it up hardest were often the ones already moving fastest: good tests, few outages, room to try things. Part of your gap is who picked the tool, not what the tool changed.
The partial fix is to match the starting points. Compare both groups on output and quality *before* they started, then report the difference net of the gap that was already there. Label it compared — a real difference between two similar teams inside your own company. A difference is not proof of a cause.
4. Dividing by the wrong number
Coding is about 14% of an engineer's day, per Microsoft research cited by DX. The rest is review, meetings, planning, debugging, waiting and switching tasks.
Run that through. Even if a tool cut coding time in half — a heroic guess — the most one person can save is about 7%. So when someone reports a 30% gain across a whole company, the math needs most of it to come from somewhere other than typing code. Not impossible. But the model has to say where, and almost none do.
DX counted a median gain of 7.76% in pull requests merged. That sits close to the same ceiling.
5. You only looked at the pilot team
You counted the pilot team. Pilot teams are volunteers with leadership attention, a keen lead, and usually the easiest code in the building. They are not your middle team.
DX's "AI Efficiency Plateau" study (May 2026) covered 400+ companies from Nov 2024 to Feb 2026, and published the whole spread instead of the average. Bottom tenth: −3%. Bottom quarter: 2%. Middle: 8%. Top quarter: 17%. Top tenth: 44%. Scale a top-tenth pilot across the whole company and you are off by about five times. Watch the bottom end too: some teams got slower, and the counts show it.
6. Asking people instead of counting
The most useful number here is a gap, not a size. METR's July 2025 randomized trial followed 16 experienced open-source developers. They believed AI made them about 20% faster, while the stopwatch showed them about 19% slower. That is a gap of roughly 39 points.
Cite that gap, and cite the walkback with it. METR's February 2026 update put real limits on the first result, because the late-2025 follow-up failed to repeat it. Returning developers came in at −18%, with a range of −38% to +9%. New ones came in at −4%, range −15% to +9%. Both ranges cross zero, and METR now reads developers as more sped up in early 2026. The size never held. Quoting "19% slower" as settled fact quotes a result its own authors have walked back.
What survives is the shape of it. What people said and what the clock said disagreed, by a lot, under test conditions. That is a problem for any return model built on a survey question. DORA's 2025 report (about 5,000 people) found over 80% believe AI made them more productive — while 30% report little or no trust in AI-written code. Stack Overflow's 2025 Developer Survey (about 49,000 people) found only 3.1% highly trust how accurate AI output is, and 45.2% say debugging AI-written code takes *more* time than writing it themselves. Belief and counting are not the same tool.
Eight separate errors that all lean the same way do not average out. They stack.
7. One number with no range
A single number with no range claims a precision you do not have. "AI delivered 24% return" says you could tell 24% from 19%, and nothing in your tracking can do that.
DORA's 2026 work shows how to publish this honestly. It reports a 39% first-year return: $11.6M of value on $8.4M spent, for a 500-person company, with payback in about 8 months. It is presented as a model, with its assumptions stated. One of them is a change failure rate rising from 5% to 6%, at a cost of $344,000. It is a model. Google says so. Present it as one or do not present it.
8. The plateau
DX's plateau data found 69.7% of developers hit peak time saved within two quarters, and 66.1% then watched that saving shrink. Counting in Q1 and multiplying by four assumes a flat line the data does not show. Only annualize once you hold the later quarters, and if you must project, project the drop.
What to report instead
Four rules. They give you a smaller number and one you can defend.
Lead with counted savings. Start with idle seats: a seat is a paid license for one person, and an idle one is a paid seat nobody has used lately. Then duplicate tools, plans bigger than you need, and extra usage charges nobody planned for. This is cash you stop spending, read straight off billing and usage systems, and bankable this quarter. It is the dullest line on the page and the only one nobody can argue with. And pricing that charges by use hides more of it than seats do.
Put estimated gains beside it, labeled and ranged. Beside, not instead of. Show the assumptions, show a band rather than a point, and test each assumption in both directions. If a 20% swing in one of them flips your answer from win to loss, that is the headline, not a footnote.
Show each speed number next to what it cost. How fast work merged, beside how much got redone within 21 days. Pull request volume, beside bugs per request and requests merged with no review. Deploy frequency, beside outages per request. If review time rose more than cycle time fell, say so on the same line, not three sections later.
Put a proof level on every line. A proof level tells the reader how much guessing sits behind a number.
Counted — read straight off a bill, a seat list or a usage report. No guessing. Two people pulling the same numbers would get the same answer.
Compared — the difference between two similar teams inside your own company, over the same period. Real signal, but a difference is not proof of a cause.
Estimated — worked out from assumptions you can see and change. Always shown as a range, never a single number.
A page where each figure carries its level is a page a CFO can defend to an audit board. In practice that is a one-page AI cost report — one page that shows what your AI tools cost and what you get back. Give each line three columns: the figure, its proof level, and the assumption or cost it rests on. The six measurement layers behind it are written up on their own, as is the case for ranges over single numbers.
The cleanest proof that any of this is needed is Harness's May 2026 survey of 700 practitioners. It found 89% of leaders say their metrics reflect AI's impact well, and 94% admit those same metrics leave out tech debt, checking time and burnout. Same people, both numbers. Confidence and accuracy have come apart, and confidence is winning.
The honest caveats
None of this gets you a clean cause-and-effect number. No AI cost report can pull AI apart from everything else that changed in the period — headcount, process, product mix, a reorg. The costs of speed also show up late: redone work lands weeks after the request that caused it, so a recent quarter always looks better than it will. Matching teams is weak when the teams are small. And the sturdiest line on the page, counted savings, is a cost cut rather than proof that anyone got more done. If you re-run it with better data and the range narrows toward zero, that is the method working, not failing.
Keep reading
Find your recoverable AI spend
Start with the line that needs no model. Seats billing without being used, duplicate tools, extra usage charges — across all your vendors at once. Snapshot shows it as a one-page AI cost report, read-only, with a proof level on every figure.
`Start free` → /login?src=blog_ai-roi-number-is-wrong
Find your recoverable AI spend
Spendassay turns this from an afternoon of spreadsheets into a live, proof-level audit with the recovery attached.