← All posts

June 18, 2026 · 7 min read · Finance operations

What an AI cost report is, and how to build one

One page. Every dollar AI costs you on one side. Every effect you can put a number on next to it. And a label on each line telling the reader how much guessing sits behind it.

By Spendassay Research

An AI cost report is one page that shows what your AI tools cost and what you get back. It is not a dashboard, and it is not a single ROI figure. It is a statement, and every line on it says how it was produced. A line can be a bill for seats — a seat is a paid license for one person. It can be a charge for tokens, the unit AI vendors bill for, roughly a few characters of text. It can be a guess at hours saved. The page tells you which.

Every number carries a proof level, so you know how much guessing sits behind it.

Counted — read straight off a bill, a seat list or a usage report. No guessing. Two people pulling the same numbers would get the same answer.

Compared — the difference between two similar teams inside your own company, over the same period. Real signal, but a difference is not proof of a cause.

Estimated — worked out from assumptions you can see and change. Always shown as a range, never a single number.

The proof levels are the whole point. Mix a counted license bill with an estimated productivity gain, print one total, and you have destroyed the thing that made the page useful. Keeping them apart lets a CFO knock down the soft lines and still act on the hard ones.

Deloitte has published a CFO guide to AI token costs, which governs this same statement from the cost side. A Big Four firm treating it that way is a credibility signal. The format is becoming standard practice, not a vendor coinage.

The cost side

Four lines, and the fourth is the one everyone leaves out.

Seats. Per-user licenses across every AI vendor, at the rate you negotiated, not at list. Counted. The easy line, and usually the only one people build.

Token and usage billing. What you pay by use: API tokens, premium requests, credits, agent runs. Counted — but only if you can tie the usage to a team, and for several vendors the usage record names a key rather than a person. Model prices move this line without anyone approving a budget. Claude Opus 4.8 lists at $5 in / $25 out per million tokens, against Sonnet 5 at $2 / $10. So a routing change is a budget change.

Infrastructure. Model endpoints, vector stores, evaluation harnesses, the middle layer every request passes through, and the tooling that watches it all. Counted, usually through cloud cost tags, and usually incomplete on the first pass.

Admin and review time. The hidden line. Someone reconciles invoices, runs access reviews, chases subscriptions nobody logged, and reviews the code the AI wrote. Work it out — hours per month times a fully loaded rate — and mark it estimated. Harness surveyed 700 practitioners in May 2026 and found 81% report developers spending more time in code review since AI, with 28% reporting a rise of more than 30%. If that time is missing, your cost side is understated by more than any discount will win back.

The value side

Three lines, in falling order of how much you should trust them. Hard proof first, so the soft proof is read in its shadow.

Recovered spend. Seats cancelled, duplicate tools merged, plans cut to the size you need, extra usage charges stopped. Counted, stated in dollars, checkable against the next invoice. This is the only value line a hostile auditor will accept without argument. In year one it is usually the largest, too.

Change in throughput. Throughput is how much work gets finished. Compared, at best: two matched teams inside your company, one using AI and one not, over a set window. Never call it counted. You cannot run a clean experiment on your own org. The public spread should keep you humble. DX's AI Efficiency Plateau study covered 400+ companies from November 2024 to February 2026. It found a median gain in pull requests shipped of 7.76%, with the bottom 10% at −3% and the top 10% at +44%. It also found 66.1% of developers saw early time savings decline after peaking within two quarters. A gain seen in month three does not stretch across a year.

Estimated value of time saved. Hours saved, turned into dollars. Always a range, with the assumptions printed beside it. Always show how far the answer moves if an assumption is wrong. Mark it estimated and expect a hostile reader to knock it to zero. That is fine. It should not be holding anything up.

Faros AI looked at 10,000+ developers across 1,255 teams and found AI users finish 21% more tasks and merge 98% more pull requests — and no significant link between AI take-up and improvement at company level. A value side built entirely from individual gains is estimating something that may never reach the income statement.

The line that shows what the speed cost you

Some quality measures move the wrong way when speed goes up. That is the cost of going faster, and it belongs beside throughput. Same page, same review. Not in an appendix.

Four belong there: rework and code churn within 21 days, bugs per pull request, review time per pull request, and incidents per pull request. Churn is code rewritten or deleted soon after it ships.

The evidence that these move is not thin. DORA 2025, n≈5,000, found AI positively related to throughput and negatively related to delivery stability. Faros found +9% bugs per developer, +154% pull request size and +91% review time. GitClear's analysis of 623 million changes found refactoring down 70% since 2022.

The gap between what teams count and what they know is the striking part. In the Harness data, 89% of leaders say their metrics accurately reflect AI's impact — while 94% admit tech debt, validation time and burnout are missing from those same metrics. Both cannot be true.

A throughput number printed next to what the speed cost you is a report. Without that, it is a pitch.

A statement is what makes a number defensible. The number alone never was.

We count each dollar once

One rule, and it stops most of the inflation in vendor ROI decks. A saved hour is counted once. If four hours of engineering time were freed, that time either produced new features or avoided a hire. It cannot be booked as both. Pick the one that actually happened, write down which, and move on.

The same rule covers recovered spend. A cancelled seat is a cost cut, not also a productivity gain. Every double count is a place where a reader finds a flaw. Once they find one, they stop trusting the whole page.

The worked layout

Rebuild this in a spreadsheet. One row per line, one column for the proof level, one for the window.

| Line | Amount | Proof level | Source | |---|---|---|---| | Cost | | | | | Seats — all vendors | $X | Counted | Vendor billing | | Token / usage billing | $X | Counted | Usage APIs, tied to teams | | AI infrastructure | $X | Counted | Cloud cost tags | | Admin + review time | $X–Y | Estimated | Hours × loaded rate | | Total cost | $X | | | | Value | | | | | Recovered spend | $X | Counted | Seats cancelled, plans cut | | Change in throughput | $X–Y | Compared | Matched team comparison | | Value of time saved | $X–Y | Estimated | Stated assumptions, ±20% | | Total value | $X–Y | | | | What the speed cost you | | | | | Rework ≤21 days | % | Counted | Delivery data | | Bugs per PR | n | Counted | Delivery data | | Review time per PR | hrs | Counted | Delivery data | | Incidents per PR | n | Counted | Delivery data | | Net | A range, not one number | | |

Cutting seats you do not use means dropping your contracted seat count or plan at renewal, down to what people actually used. It is the opposite of the clause vendors write in by default, the one that lets you add seats mid-contract.

The bottom row is a band, not a figure. Move every estimated input by ±20% and publish the range that comes out. Any tool that hands you one ROI number has hidden its assumptions rather than settled them. That is the same argument as ranges over fake precision.

Simplifying it for a smaller team

Below about 50 engineers, most of this collapses. Build three lines: total AI cost, recovered spend, and one measure of what the speed cost you that you can actually pull. Skip the team comparison. Comparing eight engineers against eight others is noise with a decimal point.

That three-line version is still an AI cost report. Costs, effects, a proof level on every line. Small and true beats large and estimated.

Why the statement matters more than the number

KPMG's Global AI Pulse, Q2 2026 asked 2,145 C-suite leaders across 20 countries. 7% report established ROI on AI. Among leaders who can see their costs clearly, 15%. Among those who cannot, 3%.

Five times the odds, from clear sight of the cost side alone. Not from better models. Just from seeing the costs well enough to set a defensible number beside them. That is what the page buys you. The number was never the deliverable.

The honest caveats

The value side is softer than the cost side and always will be. A change in throughput is a comparison between teams, not an experiment. You cannot randomise your own engineering org, and matched teams differ in ways you never controlled for. The estimated value of time saved rests on an hourly rate and an hours-saved guess, and a determined skeptic can halve either one. The quality lines tell you quality moved. They do not prove AI moved it. Build the page anyway. Its job is not certainty. Its job is to put the uncertainty where you can see it, so a reader knows which line to argue with.

Find your recoverable AI spend

The recovered-spend line is the one you can build this week. It is also the only value line that is fully counted. Spendassay produces a one-page AI cost report on the free Snapshot plan — see a filled-in example before you connect anything.

`Start free` → /login?src=blog_what-is-an-ai-cost-report

Find your recoverable AI spend

Spendassay turns this from an afternoon of spreadsheets into a live, proof-level audit with the recovery attached.

Practical, evidence-first notes on AI spend. A couple a month. No spam, unsubscribe anytime.