July 23, 2026 · 7 min read · Finance operations
How to forecast next year's AI budget
AI spend is two lines, not one. Seats — paid licenses, one per person — scale with headcount and can be capped. The meter — the part of the bill that goes up with use — scales with workload and has no natural ceiling. Forecast them as a single number and you will miss. The only question is by how much, and in which direction.
Here is the method, before the reasoning. Forecast seats from the bottom up: next year's hiring plan, times a stated coverage assumption about what share of engineers get which tool. Forecast usage spend on its own, from a trailing run-rate times a growth assumption with a named driver behind it. Publish three scenarios rather than one number. Then set tripwires that force a re-forecast mid-year, so a miss arrives as a signal instead of a shock in month eight.
Most AI budgets miss for a reason that has nothing to do with arithmetic. The two halves of the number behave nothing alike, and collapsing them into one line destroys exactly the information you need to see the miss coming.
Why the single-line AI budget misses
A seat is a decision your company makes a few times a year. Someone asks for a license, someone approves it, and it renews on a date you control. Seat spend moves in steps you can see, and it can be capped, because you can stop handing out licenses. It lags headcount by a quarter at most.
The meter is a decision someone makes every hour. Every prompt, every agent run, every long call adds to it. Nobody approves it. It responds to how deeply people use the tools, to the kind of work they do, and to which model is named in a config file. It has no ceiling except one you build.
Zylo's 2026 SaaS Management Index reports AI-native app spend up 108% year over year, and 78% of IT leaders hitting unexpected AI or usage charges. That second figure is the tell. Surprise charges on a seat-based product are rare, because you know how many seats you bought. Surprise charges come from usage billing, and 78% is the base rate, not an outlier.
So the meter gets its own budget line, its own forecast method, and its own named owner. If those three things are not true in your plan, you have no forecast for the half of the spend that actually moves.
Forecast seats bottom-up from the hiring plan
Start with the hiring plan by month, not the year-end headcount. A December start date adds one month of seat cost, not twelve. Companies that forecast off the year-end number routinely overstate seat spend by 20–30%.
Then apply a coverage assumption per tool, stated as a share of a group you have defined. Not "everyone gets AI tools." Something you can check later: 100% of backend and frontend engineers get an inline assistant, 30% of engineers get an agent tool, 0% of data analysts get either until Q3. Coverage is the assumption most likely to be wrong and the easiest to check against reality later, which makes it the most valuable one to write down.
Multiply coverage by the seat price you actually pay, including the lines that never appear on the pricing page. GitHub Copilot Enterprise lists at $39/user, but it requires GitHub Enterprise Cloud at roughly $21/user, so the real seat is about $60 (prices per GitHub's plans page, read July 2026). Budgets built on the sticker rather than the whole stack miss by 50% on that line alone.
Finally, subtract what you already have the right to remove. Most seat forecasts roll last year's license count forward, licenses nobody ever opened included. The proof that makes idle seats visible — a paid seat nobody has used lately — is the same proof that makes them removable at renewal.
Seats are a decision you make a few times a year. The meter is a decision someone makes every hour.
Forecast the meter from run-rate and a named driver
The meter cannot be forecast from the top down off a headcount number. Tokens — the unit AI vendors bill for, roughly a few characters of text — vary per engineer by more than tenfold across usage profiles. Forecast the meter from your own trailing run-rate instead.
Take the last three months of usage spend and put it in the same format. Per active engineer per working day is the most stable thing to divide by, because it strips out both headcount growth and the December problem.
Then apply a growth assumption, and name the driver. "20% growth" is not a forecast, it is a mood. This is a forecast: the meter grows 40% because agent coding moves from 30% to 60% of engineers, and agent tasks burn roughly ten times the tokens that inline completion does. When it turns out wrong, you know which clause broke.
Three drivers move the meter, and they compound.
- Breadth. More engineers using metered tools. Roughly linear, and the easiest to predict, because it follows your coverage assumption.
- Depth. The same engineers using them harder. This is where the tail lives. DX's $200–600 per engineer per month is mostly a depth spread, not a pricing difference — and it describes other teams, so track it in your own run-rate rather than adopt it as a target.
- Model mix. The one finance usually treats as a technical detail. Claude Opus 4.8 runs $5 in / $25 out per million tokens; Claude Sonnet 5 runs $2 in / $10 out (read July 23, 2026). A full swap in either direction is a 60% swing on identical token volume. Nobody files a change request to move a default model. Your bill moves anyway.
Note that model prices have moved in both directions over the past two years. Do not build a forecast that assumes only decline, and do not build one that assumes only increase. Build one that states which prices you assumed and the date you read them.
Publish three scenarios, not one number
One number invites a false debate about whether it is correct. Three scenarios move the conversation to which assumptions you are willing to defend, which is the conversation worth having.
| Scenario | Seat assumption | Meter assumption | What it is for | |---|---|---|---| | Low | Coverage flat, hiring comes in 20% under plan | Run-rate grows with headcount only | The floor you commit to | | Base | Coverage as planned, hiring plan hits | Run-rate per engineer grows 30–50% on a named driver | The number in the plan | | High | Coverage expands to a second group | Agent shift lands early; model mix moves up-tier | The reserve you ask for |
The spread between Low and High is the honest output. If it is narrow, your spend is mostly seats and you have a buying problem. If it is wide, your spend is mostly meter and you have a control problem. Confusing one for the other is how companies end up negotiating hard on a line that was never moving.
The things that break the forecast
Four failures account for most of the miss, and each has a specific tell.
1. A model price change. Vendors reprice per-token rates with little notice. Watch the in and out rates on the models your default config actually calls.
2. A shift to agent work. Tokens per task can jump tenfold when a team adopts long-running agents. It shows up as a step change in tokens per active developer, not a gradual slope.
3. A vendor repricing or plan change mid-term. Included-usage allowances get re-cut, credits get redefined, tiers get renamed. Your cost per unit changes while your usage stays flat. The mechanics are covered in the traps hiding in AI tool pricing.
4. Take-up above or below plan. Under-use looks like a win in month three and a write-off at renewal. Over-use looks like a crisis, and it is often the cheapest thing in the plan.
The renewal calendar is a forecast input
Pull every AI contract into one calendar with four columns. Renewal date. Committed seat count. Seats actually used. And your right to cut seats you do not use — the contract right to reduce the committed quantity at renewal without penalty. Most teams find they hold that right and never used it, because nobody assembled the usage proof before the auto-renew window closed.
The calendar tells you which quarters give you room to move, and when to start building the case. Proof assembled the week before a renewal is a request. Proof assembled a quarter before is a negotiation.
Tripwires that force a re-forecast
Set thresholds now, while nobody is defensive, and route them to a named person.
- Usage-based spend runs over 115% of the monthly budget for two months in a row.
- Tokens per active developer per day rise more than 40% month over month.
- Any vendor announces a price or plan change on a tool you hold seats in.
- Seat use on any vendor drops below 70%.
- A default model changes in any shared config.
Each one triggers a re-forecast, not an escalation. The point is not to catch someone out. It is to turn a year-end shock into a mid-quarter adjustment.
The honest caveats
This method will not make your forecast right. A trailing run-rate is a weak predictor when a team is mid-adoption, which is where most engineering teams sit right now, and a three-month base rate taken during a step change will mislead you in both directions. The DX figures cited here describe other companies' usage, not yours, and span a 3x range on their own. Treat them as range-setting, not as benchmarks to plan against. Vendor pricing pages were read in July 2026 and change without notice. And a scenario band is only as honest as the assumptions written beside it. The goal is not a forecast that is right. It is a forecast whose assumptions are named, so when it is wrong you know which one broke.
Keep reading
Find your recoverable AI spend
Spendassay reads seats and token usage across your vendors and splits them into the two lines your forecast needs, including a per-active-developer meter run-rate you can carry straight into the plan. Snapshot returns a one-page AI cost report — one page that shows what your AI tools cost and what you get back — free, read-only, no card.
`Start free` → /login?src=blog_forecast-next-year-ai-budget
Find your recoverable AI spend
Spendassay turns this from an afternoon of spreadsheets into a live, proof-level audit with the recovery attached.