AI cost · Google
Vertex AI cost: what you actually pay
Gemini and hosted third-party models (Claude, Llama) price at roughly direct-API parity; the real spend drivers are the price doubling above 200K context, 'thinking' tokens billed as output, and per-hour context-cache storage charged whether or not the cache is reused.
Public pricing
| Plan | Price | Billing | Includes |
|---|---|---|---|
| Pay-as-you-go | CustomStandard on-demand token pricing across Gemini and Model Garden models. | usage | — |
| Provisioned Throughput - 1 month | $2,700 / per GSU/mo (Global)GSU = Generative AI Scale Unit; 1-week is ~$1,200/GSU, 3-month ~$2,400/GSU/mo, 1-year ~$2,000/GSU/mo. | custom | — |
Token rates (USD per 1M tokens)
| Model | Input | Output | Cache read |
|---|---|---|---|
| Gemini 2.5 ProRate for prompts <=200K tokens context; >200K context roughly doubles to $2.50 in / $15.00 out. | $1.25 | $10 | — |
| Gemini 2.5 FlashOutput price includes 'thinking' tokens by default. | $0.3 | $2.5 | $0.03 |
| Claude Sonnet 4.6 (Anthropic, via Model Garden)Parity with direct Anthropic API/Bedrock pricing; regional (non-global) endpoints are reported to carry a premium over global pricing. | $3 | $15 | — |
Gemini and hosted third-party models (Claude, Llama) price at parity with their native APIs. Vertex's own twist: Gemini Pro-tier context pricing roughly doubles above 200K tokens, and 'thinking'/reasoning tokens are billed as output tokens by default. Context caching bills separately by storage-hour ($1-4.50 per MTok/hour depending on model tier) whether or not the cache is ever reused. Provisioned Throughput is sold in Generative AI Scale Units (GSUs) on terms from 1 week to 1 year, with longer commitments earning a lower per-GSU rate. As of mid-2026 the catalog spans both the newer Gemini 3.x generation and still-available Gemini 2.5 models.
Last updated: 2026-07-22. Sourced from Vertex AI Generative AI Pricing (official) ↗, Google Vertex AI Pricing breakdown (CloudZero) ↗.
Where the real Vertex AI cost hides
The sticker is the smallest part of the story. For Vertex AI, these are the line items that quietly inflate the bill:
- Context above 200K tokens roughly doubles the input (and often output) price on Gemini Pro-tier models — long-document/RAG workloads that creep past the threshold get an invisible rate change mid-conversation.
- 'Thinking'/reasoning tokens are billed as output tokens even when never shown to the user — reasoning-heavy tasks can cost far more than the visible response length suggests.
- Claude and Llama models on Vertex AI Model Garden bill at the same per-token rate as calling the provider directly, plus a reported premium on regional (non-global) endpoints chosen for latency or compliance reasons.
- Context caching bills storage by the hour ($1-4.50/MTok/hour) regardless of whether the cache is ever hit again — caching long documents 'just in case' can cost more than simply re-sending them.
- Provisioned Throughput (GSUs) is quoted as a flat per-unit commitment from 1 week to 1 year — shorter terms carry the highest per-GSU rate, penalizing teams still validating volume.
What an audit finds on Vertex AI
Apps default to Gemini Pro (or Claude Opus/Sonnet via Model Garden) for every call, let prompts drift past the 200K context-doubling threshold, and leave 'thinking' budgets uncapped — routing to Flash/Flash-Lite, trimming context, and capping thinking tokens routinely cuts spend 70-90% with little quality loss for most tasks.
Vertex AI pricing FAQ
Does Vertex AI cost more than calling Gemini or Claude directly?
No — Gemini and Claude on Vertex price at parity with their direct APIs (Claude reportedly carries a premium only on non-global regional endpoints); Vertex's overhead comes from GCP networking, storage, and Search-grounding fees around the call.
What happens if my prompt exceeds 200K tokens?
Gemini Pro-tier models roughly double both input and output price above 200K tokens of context; Flash and Flash-Lite tiers are not subject to this threshold at the same rates.
Is Provisioned Throughput worth it?
Only at sustained, predictable volume — GSU commitments run 1 week to 1 year, and shorter terms cost more per unit, so it doesn't help bursty or early-stage workloads.
What is Vertex AI really costing you?
Spendassay connects your Vertex AI spend to actual usage and engineering output — a proof-level AI cost report with the wasted dollars named. Neutral across every AI vendor.
Free · read-only · no card
Other ai infrastructure pricing
More: the complete AI cost management guide · every tool, with prices · cost per outcome · compare platforms · how we measure