An independent auditor is only useful if you can verify it. This page shows how every number in a Spendassay audit gets built. You will seeThe Six-Layer Audit, the Proof Levels scale behind every figure, the cost of going faster we refuse to hide, and an ROI model that reports a range. The audit runs continuously. It recomputes every day and alerts you when the numbers drift.
By Spendassay Research
No single metric survives a skeptical CFO. We call this The Six-Layer Audit: value builds up in six layers, each auditable on its own. The higher the layer, the more inference it carries, and the weaker the proof level we're willing to stamp on it. Each rung below carries the level it earns.
Seats provisioned, active vs. dormant, license and API-key inventory across every vendor. Straight from billing and admin APIs. The denominator for everything above it.
Suggestions accepted, tokens consumed, sessions, feature mix. Read from vendor usage data. A signal of engagement, not yet a claim of value.
PR throughput, cycle time, review latency, deploy frequency. Team-level, from Git and CI metadata. What you'd track with or without AI.
Rework, bugs per PR, unreviewed merges, incidents. Rendered beside the throughput numbers, never behind them. This layer cannot be turned off.
Delivery and quality deltas priced with loaded engineering cost, set against real AI spend. Measured against a matched low-AI cohort inside your company wherever possible, not an industry constant.
Net value as a band with a payback window. Every input editable, every formula visible, always labeled as what it is: a model.
The scale is called Proof Levels, and every number carries one of its three levels, telling you exactly how much inference sits between the raw log and the figure on screen. A lower level never dresses up as a higher one. The full definitions.
Read directly from a system of record: billing, vendor usage data, Git, CI, the incident tracker. No modeling, no estimation. Two people pulling the same window get the same number.
A measured difference between comparable groups: an AI-heavy team against a matched low-AI cohort, same company, same period, same work profile. Stronger than any industry benchmark, but it still attributes a gap to AI, so it sits one rung down.
A projection built on explicit, editable assumptions: dollar values, adoption curves, discount rates. Useful for planning, honest about its uncertainty. Always ships with a sensitivity range and never appears without its label.
These aren't settings a customer can toggle. They're architectural, which is why the audit survives a works-council review and a security questionnaire.
Reporting is team-level, always. There is no per-person ranking, no leaderboard, and no way to slice a metric down to a single engineer. Data that can be turned against a person is data we refuse to build.
We read commit and PR statistics: counts, timestamps, churn, review events. We never read, store, or transmit diffs, source files, or prompts. The measurement only needs metadata, so metadata is all we take.
The audit reads through read-scoped access. It cannot merge a PR, change a license, or revoke a seat: an auditor that can change the books isn't an auditor. Acting on a finding is a separate, opt-in step on the Control plan (private beta), and it runs through your own provider keys.
No cohort smaller than eight people renders a productivity figure. Below the floor, the number is suppressed rather than risk re-identifying an individual from an aggregate.
A claim is only checkable if its terms are. These are the thresholds the engine applies, not a description of them.
A seat is unused when the license has had no recorded activity for 60 days and is still being billed.
A seat is underused when it shows activity on fewer than 3 days in a month: opened, but not enough to earn its price.
At a renewal we apply a stricter test than the 60-day bar, because the honest question in a negotiation is who used the tool in the last 30 days.
Seat and subscription spend is the recurring, per-license part of the AI bill: the seats you are billed for on a fixed plan, plus flat subscription fees. It excludes metered usage (tokens, agent runs, on-demand compute), which is not recoverable by cancelling a seat, and it excludes infrastructure.
It is the denominator of the 10 to 15% band we publish. We scope it deliberately: a token-heavy stack recovers less on seats, so a band stated against the whole AI bill would be true only for the stacks that happen to look like our demo.
The easiest way to lie with an AI dashboard is to show acceleration and stop. Spendassay renders every throughput gain beside the cost of going faster, on the same surface, at the same time.
Lines rewritten or deleted within three weeks of merge. Code that never stabilizes is negative throughput.
Defects traced back to a change. Shipping faster only counts if the defect rate holds.
Share of PRs merged without substantive human review. Velocity bought by skipping review is debt, not delivery.
Production incidents attributable to a change. The tax on acceleration that spikes first.
In Faros AI's 2026 analysis of 22,000 developers, teams that pushed epic throughput up with AI saw incidents per PR rise roughly 3x while code churn climbed over the same window. The velocity chart alone looked like a win. That gap is why the cost of going faster is product law here, not a preference.
Layer 6 is the only place Spendassay models, and it models out loud: a band with a payback window, inputs you can edit, and the formula printed next to the result.
Loaded cost per engineer, cost per incident, adoption ramp: all inputs, all editable. Change one and the range recomputes in front of you. Defaults ship visible.
The headline is a band, not a point. The envelope keeps the model's uncertainty legible instead of laundering it into false precision.
A saved hour is counted once. It cannot be booked as features shipped and as headcount avoided. Double-counting is how AI ROI decks reach implausible numbers.
Adoption dips before it pays back, as teams learn to review machine-drafted code. We model the dip and place typical payback at 6–18 months, not day one.
Spendassay doesn't ask you to trust a proprietary productivity score. The method is assembled from public research on AI and software delivery, weighted toward the studies that constrain us rather than flatter us.
The perception gap is the point: developers were about 39 percentage points off about their own AI speedup — they estimated a 20% gain on real tasks while the original measurement found a 19% slowdown. A follow-up did not reproduce that slowdown (METR, 24 Feb 2026 update): returning developers came in at -18% (range -38% to +9%) and new developers at -4% (range -15% to +9%), both ranges crossing zero, and METR now says developers are more sped up from AI tools in early 2026. Why we report ranges and never trust self-report.
Teams accelerating epic throughput with AI saw incidents per PR rise roughly 3x. The finding that makes the cost of going faster non-negotiable.
Large-corpus evidence of rising churn and copy-pasted code alongside assistant adoption. Grounds our churn and duplication metrics.
AI amplifies the system it lands in: strong delivery gets faster, weak delivery breaks faster. Why we compare against your existing DORA baseline.
Copilot acceptance rate tracks perceived productivity, not output. Why acceptance stays a coverage signal in Layers 1–2, never a delivery outcome.
Documents the adoption J-curve and the verification tax. Source of our 6–18 month payback window and the single-counting rule.
Minimum-cohort floors for reporting team analytics without re-identifying individuals. Basis for our hard 8-person floor.
A multi-dimensional measurement frame that refuses to collapse AI value into one number. Shapes the layered stack itself.
Every number in a Spendassay audit carries its proof level, every throughput chart the cost of going faster, and every ROI figure is a range you can re-derive.