The framework
What is a proof level?
A proof level states how much guessing sits behind a number. There are three, and every figure Spendassay publishes carries exactly one: Counted is read straight from a system of record, Compared is a measured difference between similar teams inside one company, and Estimated is a projection from assumptions you can see and change.
The scale exists because the most common way to mislead with a dashboard is not to state something false. It is to present a modeled number in the same typeface as a counted one and let the reader assume they are the same kind of thing.
The three levels
Counted, Compared, Estimated
Counted
A number you can re-derive from a system of record.
Counted is read straight off a bill, a seat list or a usage report. No guessing. Two people pulling the same numbers would get the same answer.
Seat counts, invoice totals, token volumes, last-activity dates. If two people pull it independently they get the same answer, and the vendor whose console it came from can check it. Recovery findings live here, which is why they are the part of an audit you can act on the same week.
Compared
A measured difference between comparable groups, inside one company.
Compared is the difference between two similar teams inside your own company, over the same period. Real signal, but a difference is not proof of a cause.
Two similar teams, same period, same kind of work, one with heavy AI adoption and one without. The comparison controls for the things a cross-company benchmark cannot. It is real signal and it is not proof of a cause, because teams differ in ways no dataset fully captures.
Estimated
A projection from assumptions you can see and change.
Estimated is worked out from assumptions you can see and change. Always shown as a range, never a single number.
Anything that requires a rate, a price of an engineering hour, or a forecast. Always rendered as a range with a sensitivity band, never as a point. A modeled number that arrives as a single confident figure has been laundered into looking like a counted one.
What keeps it honest
Four rules, or the labels are decoration
A three-word scale is easy to adopt and easy to render meaningless. These are the rules that make the label load-bearing, and they are enforced in code rather than in a style guide.
A lower level never dresses up as a higher one.
The level is attached to the number at the point it is computed, not chosen at render time by whoever is presenting. An estimate cannot be promoted by putting it in a bigger font.
Mixed inputs inherit the weakest level.
If a figure combines a counted seat price with a modeled adoption rate, the result is Estimated. Averaging the confidence of the inputs would overstate the confidence of the output.
Every number opens its formula.
A level is a claim about provenance, so provenance has to be inspectable. Each figure exposes the arithmetic and the rows underneath it, and a level with no drill-through is a label rather than a proof.
The headline is Counted, or it says why not.
The first number a reader meets should be the one that needs no argument. When a headline must be modeled, it carries its range and its label in the same line, not in a footnote.
The full measurement stack these levels are assigned from is on the methodology page, and you can see the levels attached to real findings in the sample report.
Common questions
Using the scale
- What is a proof level?
- A three-level scale stating how much inference sits between a number and the system it came from. Counted is read directly from a system of record. Compared is a measured difference between similar groups in the same company. Estimated is a projection from stated assumptions, always shown as a range.
- Why not just say confidence, or high, medium, low?
- Because those describe a feeling about a number rather than its origin. A confidence score can be asserted; a proof level can be checked. Counted means you can go and re-derive it, which is a claim someone can prove wrong.
- Which proof level do recovery findings carry?
- Counted. Idle seats, tool overlap and renewal exposure are read from seat rosters, invoices and usage APIs, so they can be verified against the vendor's own console. That is why the recoverable band, 10 to 15% of seat and subscription spend, is quoted separately from any productivity estimate.
- Can a number move between levels?
- Yes, and it should. An estimate becomes Compared once there is enough history to compare matched teams, and a Compared figure never becomes Counted, because a difference between groups is not a reading from a system of record. Movement in the other direction, a number quietly gaining confidence it did not earn, is the failure this scale exists to prevent.
- Is this specific to AI spend?
- The scale is general. It matters most where a vendor benefits from the ambiguity, which is exactly the position an AI tool vendor is in when reporting on its own impact. An independent auditor should have to say, on every figure, how much of it is a reading and how much is a model.
See the levels on your own numbers
Connect read-only sources and every figure arrives with its level and its formula attached. Free to start, about 10 minutes.
Related: cost per outcome · the AI cost management guide · glossary · what an AI spend audit is