AI cost · Google
Vertex AI cost: what you actually pay
Gemini and hosted third-party models (Claude, Llama) price at roughly direct-API parity; the real spend drivers are the price doubling above 200K context, 'thinking' tokens billed as output, and per-hour context-cache storage charged whether or not the cache is reused.
Public pricing
| Plan | Price | Billing | Includes |
|---|---|---|---|
| Pay-as-you-go | CustomStandard on-demand token pricing across Gemini and Model Garden models. | usage | — |
| Provisioned Throughput - 1 month | $2,700 / per GSU/mo (Global)GSU = Generative AI Scale Unit; 1-week is ~$1,200/GSU, 3-month ~$2,400/GSU/mo, 1-year ~$2,000/GSU/mo. | custom | — |
Token rates (USD per 1M tokens)
| Model | Input | Output | Cache read |
|---|---|---|---|
| Gemini 2.5 ProRate for prompts <=200K tokens context; >200K context roughly doubles to $2.50 in / $15.00 out. | $1.25 | $10 | — |
| Gemini 2.5 FlashOutput price includes 'thinking' tokens by default. | $0.3 | $2.5 | $0.03 |
| Claude Sonnet 4.6 (Anthropic, via Model Garden)Parity with direct Anthropic API/Bedrock pricing; regional (non-global) endpoints are reported to carry a premium over global pricing. | $3 | $15 | — |
Gemini and hosted third-party models (Claude, Llama) price at parity with their native APIs. Vertex's own twist: Gemini Pro-tier context pricing roughly doubles above 200K tokens, and 'thinking'/reasoning tokens are billed as output tokens by default. Context caching bills separately by storage-hour ($1-4.50 per MTok/hour depending on model tier) whether or not the cache is ever reused. Provisioned Throughput is sold in Generative AI Scale Units (GSUs) on terms from 1 week to 1 year, with longer commitments earning a lower per-GSU rate. As of mid-2026 the catalog spans both the newer Gemini 3.x generation and still-available Gemini 2.5 models.
Last updated: 2026-07-22. Sourced from Vertex AI Generative AI Pricing (official) ↗, Google Vertex AI Pricing breakdown (CloudZero) ↗.
Where the real Vertex AI cost hides
The sticker is the smallest part of the story. For Vertex AI, these are the line items that quietly inflate the bill:
- Context above 200K tokens roughly doubles the input (and often output) price on Gemini Pro-tier models — long-document/RAG workloads that creep past the threshold get an invisible rate change mid-conversation.
- 'Thinking'/reasoning tokens are billed as output tokens even when never shown to the user — reasoning-heavy tasks can cost far more than the visible response length suggests.
- Claude and Llama models on Vertex AI Model Garden bill at the same per-token rate as calling the provider directly, plus a reported premium on regional (non-global) endpoints chosen for latency or compliance reasons.
- Context caching bills storage by the hour ($1-4.50/MTok/hour) regardless of whether the cache is ever hit again — caching long documents 'just in case' can cost more than simply re-sending them.
- Provisioned Throughput (GSUs) is quoted as a flat per-unit commitment from 1 week to 1 year — shorter terms carry the highest per-GSU rate, penalizing teams still validating volume.
What an audit finds on Vertex AI
Apps default to Gemini Pro (or Claude Opus/Sonnet via Model Garden) for every call, let prompts drift past the 200K context-doubling threshold, and leave 'thinking' budgets uncapped — routing to Flash/Flash-Lite, trimming context, and capping thinking tokens routinely cuts spend 70-90% with little quality loss for most tasks.
Worked example: Vertex AI at production volume
Take 50M input and 10M output tokens a month — a modest production workload, and a round number we picked — on Gemini 2.5 Pro at the published rates above.
- Monthly
- $163
- Annual
- $1,950
- On Gemini 2.5 Flash
- $40
Same volume, $1,470 a year apart. Most workloads are a mix — some of it genuinely needs the frontier model and some of it does not — so the real question is what share of your calls is the routine kind, and that is a measurement, not a guess.
Arithmetic on the published prices above. Your mix will differ — run your own numbers.
Vertex AI pricing FAQ
Does Vertex AI cost more than calling Gemini or Claude directly?
No — Gemini and Claude on Vertex price at parity with their direct APIs (Claude reportedly carries a premium only on non-global regional endpoints); Vertex's overhead comes from GCP networking, storage, and Search-grounding fees around the call.
What happens if my prompt exceeds 200K tokens?
Gemini Pro-tier models roughly double both input and output price above 200K tokens of context; Flash and Flash-Lite tiers are not subject to this threshold at the same rates.
Is Provisioned Throughput worth it?
Only at sustained, predictable volume — GSU commitments run 1 week to 1 year, and shorter terms cost more per unit, so it doesn't help bursty or early-stage workloads.
How much can routing cut a Vertex AI bill?
On 50M input and 10M output tokens a month, Gemini 2.5 Pro costs $163 at published rates and Gemini 2.5 Flash costs $40 — $1,470 a year apart for the same volume. What you can actually move depends on how much of your traffic is routine, which is a measurement rather than a guess.
Are these Vertex AI prices current?
They were last checked against Vertex AI Generative AI Pricing (official) and Google Vertex AI Pricing breakdown (CloudZero) on 2026-07-22. Every price on this page carries that date and a link to its source, and a scheduled check fails our build when any figure goes more than 30 days unverified. Vendors do change pricing between checks — the source link is there so you can confirm before you sign.
What is Vertex AI really costing you?
Spendassay connects your Vertex AI spend to actual usage and engineering output — a proof-level AI cost report with the wasted dollars named. Neutral across every AI vendor.
Free · read-only · no card
Other ai infrastructure pricing
More: the complete AI cost management guide · every tool, with prices · cost per outcome · compare platforms · how we measure