June 16, 2026 · 6 min read · Measuring the payoff
Honest AI ROI modeling: ranges, not fake precision
A speed number on its own is propaganda. You have to show what the speed cost you. Here is how to model AI ROI a CFO can actually sign.
Three rules produce an AI ROI figure you can defend. Show every gain next to the number that catches its cost. Put a proof level on every figure. Lead with what you counted, not what you estimated — and show the estimated part as a range wide enough to be honest.
Follow them and you will publish a smaller number than your vendor's case study. You will also still have that number after the audit committee reads it. KPMG's Q2 2026 Global AI Pulse, across 2,145 C-suite leaders, found only 7% report established ROI on AI — 15% among those with strong cost visibility, 3% among those without. That gap is not about analytical talent. It is about whether the inputs could be trusted.
This is the constructive post. Why your AI ROI number is wrong lists the eight ways these models break. What the research actually says about AI coding ROI surveys the studies.
Rule 1: always show what the speed cost you
Buy speed and something else moves the wrong way. Report how fast work merged, without reporting how much of it got rewritten, and you have not shown half the picture. You have shown the flattering half, and you chose it.
The evidence here is specific. Faros AI's analysis of 10,000+ developers across 1,255 teams found two things at once. Developers using AI completed 21% more tasks and merged 98% more pull requests. They also logged 9% more bugs per developer, 154% larger PRs, and 91% more PR review time — and Faros found no significant correlation between AI adoption and company-level improvement.
The mechanism shows up in those numbers. The work never disappeared. It moved from the author to the reviewer, where nobody was counting it. A speed number reported on its own books the author's gain and expenses none of the reviewer's cost.
Pair them on the same line, not three sections apart:
- Cycle time next to work redone within 21 days
- Pull request volume next to bugs per pull request and merges nobody reviewed
- Deploy frequency next to change failure rate and incidents per pull request
If review time rose more than cycle time fell, that belongs in the headline.
Rule 2: put a proof level on every figure
Some numbers come straight off an invoice — a seat count, say, and a seat is a paid license for one person. Others come out of a model. A reader cannot tell which is which unless you say so.
Every number we show carries a proof level, so you know how much guessing sits behind it.
Counted — read straight off a bill, a seat list or a usage report. No guessing. Two people pulling the same numbers would get the same answer.
Compared — the difference between two similar teams inside your own company, over the same period. Real signal, but a difference is not proof of a cause.
Estimated — worked out from assumptions you can see and change. Always shown as a range, never a single number.
Every line on the page carries one of those three. Most AI ROI figures in circulation are estimated and presented as counted. That is the problem in one sentence, and the label is the fix.
Rule 3: estimated figures are ranges, and you show the arithmetic
A single number claims a precision your tracking cannot support. "AI delivered 24% ROI" says you can tell 24% apart from 19%. You cannot.
The width of the range is not a matter of taste. DX's "AI Efficiency Plateau" (May 2026, 400+ companies) published the full spread of throughput gain — throughput being how much work actually gets finished — rather than an average. The median was 7.76%, with P10 −3%, P25 2%, P75 17% and P90 44%. Some teams got slower. A model built on one lift figure picks a point on that curve and hides the pick.
A worked range, with the assumptions shown
Two hundred engineers, fully loaded at $200,000 each: $40M of engineering cost.
Assumption 1 (soft): coding is roughly 14% of a developer's day, per Microsoft research cited by DX. Cost you can pin to coding: $5.6M.
Assumption 2 (bounded by published data): a throughput lift of 2% to 17%, DX's P25–P75 range. Gross value:
| Lift | Gross annual value | |---|---| | 2% (P25) | $112,000 | | 8% (median) | $448,000 | | 17% (P75) | $952,000 |
Assumption 3 (the softest input here): what the speed cost you — the extra review and the rework, meaning work that has to be redone — eats 15% to 35% of the gross gain. That is an assumption, not a count. Faros's 91% jump in review time says it is not zero. Nothing published says what it is for you.
Net band: $73,000 to $809,000, against AI tool spend of about $104,000 a year for that company. ROI runs from −30% to +675%. The band crosses zero.
A band that crosses zero is not a failed model. It is an honest one — and it tells a CFO exactly which assumption to attack.
That is the useful output. At the bottom of the spread, with a heavy cost of going faster, this does not pay back. So go and count your own position on that curve. Two inputs decide the answer, and you can count both. Move the coding-share assumption 20% either way and the whole band moves with it.
Lead with what you counted
Both halves belong on the page. Only one of them is bankable.
Recoverable spend is counted, and it is immediate. Idle seats — paid seats nobody has used lately — plus tools that duplicate each other, plans bigger than you need, and usage charges nobody forecast. For the same 200-engineer company, an idle-seat audit typically returns $10,000 to $22,000 a year, read straight off the billing system. It is the least exciting line on the page and the only one nobody can argue with.
Productivity is real, and it is estimated. The gains exist. What does not exist is a defensible way to state them as one number.
So lead with $22,000 counted, and put $73,000–$809,000 estimated beside it, labeled, with its assumptions in view. A CFO can sign that page. Nobody can sign "AI generated $1.4M in value."
And count each dollar once. A saved hour is one saved hour. Book it as headcount you avoided and you cannot also book it as features shipped faster.
How to handle DORA's 39%
DORA's 2026 work reports a 39% first-year return — $11.6M of value on $8.4M invested for a 500-person company, or roughly an eight-month payback. It gets quoted constantly as a measurement. It is not one, and DORA says so. It is a model with stated assumptions, including change failure rate rising from 5% to 6% at a cost of $344,000.
That is admirably transparent, and transparency does not promote it to a higher proof level. Cite it as a model with its assumptions attached, or not at all. Then treat it as a template, because Google published the number that costs its own case money.
Why all this matters. Harness's May 2026 survey of 700 practitioners found 89% of leaders say their metrics accurately reflect AI's impact. In the same survey, 94% admit those same metrics omit technical debt, validation time and burnout. Same people, both numbers. Confidence has come apart from accuracy, and confidence is winning.
The honest caveats
None of this produces a causal number. No model separates AI from everything else that changed in the period: headcount, process, product mix, a reorg. The worked example is illustrative arithmetic on published inputs, not a prediction about your company. The assumption about what the speed cost you is the widest source of error in it.
The costs of going faster show up late, so a recent quarter always looks better than it will once the rework lands. Comparing two teams is weak when the teams are small. And the sturdiest line on the page, counted recovery, is a cost cut rather than evidence that anyone got more productive. If better data narrows your range toward zero, that is the method working.
Keep reading
Find your recoverable AI spend
Start with the half that needs no model. Spendassay builds you an AI cost report — one page that shows what your AI tools cost and what you get back. Every figure carries a proof level, and the ROI line is a range rather than a single number.
`Start free` → /login?src=blog_honest-ai-roi-modeling
Find your recoverable AI spend
Spendassay turns this from an afternoon of spreadsheets into a live, proof-level audit with the recovery attached.