AI cost · Cohere
Cohere API cost: what you actually pay
Cohere's flagship Command A (and its Reasoning/Vision/Translate variants) needs a sales conversation for real production rate limits, so budgets built on trial-key or legacy Command R pricing understate what production actually costs.
Public pricing
Token rates (USD per 1M tokens)
| Model | Input | Output | Cache read |
|---|---|---|---|
| Command Acurrent flagship; full production access gated behind contacting sales, trial keys capped at 20 req/min | $2.5 | $10 | — |
| Command R (08-2024)mid-tier, self-serve | $0.15 | $0.6 | — |
| Command R7Bcheapest current-gen model | $0.0375 | $0.15 | — |
Command A / Command A Reasoning / Vision / Translate require a sales contact for production limits (trial keys: 20 req/min, evaluation only); Command R/R7B remain self-serve at listed rates. Embed and Rerank are billed separately per call, not part of chat token pricing. Model Vault (dedicated hosting) is billed hourly/monthly ($4-$10/hr) instead of per-token.
Last updated: 2026-07-22. Sourced from Cohere Pricing (official) ↗, Command A pricing — OpenRouter ↗.
Where the real Cohere API cost hides
The sticker is the smallest part of the story. For Cohere API, these are the line items that quietly inflate the bill:
- Newest Command A Reasoning/Vision/Translate variants require contacting sales for production access; trial keys are capped at 20 requests/minute, so teams prototyping against them hit a wall before launch.
- Model Vault (dedicated deployment) bills per-hour ($4-$10/hr) or per-month regardless of utilization, so a lightly used dedicated endpoint can cost far more than serverless per-token pricing.
- Embed and Rerank are billed as separate per-token/per-call add-ons from generation — easy to under-forecast a RAG pipeline's total bill by pricing only the chat model.
What an audit finds on Cohere API
Teams budget using Command R7B's list price but route real RAG/agent traffic through Command A for quality, quietly increasing effective $/token by roughly 65x without updating cost forecasts.
Cohere API pricing FAQ
Is Cohere's cheapest model good enough for production?
Command R7B is about 65x cheaper than Command A per output token, but most teams needing strong reasoning or tool-use end up on Command A, which requires a sales conversation for full production rate limits.
Does Cohere charge for embeddings and rerank separately from chat?
Yes — Embed and Rerank are billed independently from the generation model, so a full RAG stack's real cost is generation + embed + rerank combined, not just the chat rate.
What is Cohere API really costing you?
Spendassay connects your Cohere API spend to actual usage and engineering output — a proof-level AI cost report with the wasted dollars named. Neutral across every AI vendor.
Free · read-only · no card
Other api platforms pricing
More: the complete AI cost management guide · every tool, with prices · cost per outcome · compare platforms · how we measure