AI cost · Cohere

Cohere API cost: what you actually pay

Cohere's flagship Command A (and its Reasoning/Vision/Translate variants) needs a sales conversation for real production rate limits, so budgets built on trial-key or legacy Command R pricing understate what production actually costs.

Public pricing

Token rates (USD per 1M tokens)

ModelInputOutputCache read
Command Acurrent flagship; full production access gated behind contacting sales, trial keys capped at 20 req/min$2.5$10
Command R (08-2024)mid-tier, self-serve$0.15$0.6
Command R7Bcheapest current-gen model$0.0375$0.15

Command A / Command A Reasoning / Vision / Translate require a sales contact for production limits (trial keys: 20 req/min, evaluation only); Command R/R7B remain self-serve at listed rates. Embed and Rerank are billed separately per call, not part of chat token pricing. Model Vault (dedicated hosting) is billed hourly/monthly ($4-$10/hr) instead of per-token.

Last updated: 2026-07-22. Sourced from Cohere Pricing (official), Command A pricing — OpenRouter.

Where the real Cohere API cost hides

The sticker is the smallest part of the story. For Cohere API, these are the line items that quietly inflate the bill:

  • Newest Command A Reasoning/Vision/Translate variants require contacting sales for production access; trial keys are capped at 20 requests/minute, so teams prototyping against them hit a wall before launch.
  • Model Vault (dedicated deployment) bills per-hour ($4-$10/hr) or per-month regardless of utilization, so a lightly used dedicated endpoint can cost far more than serverless per-token pricing.
  • Embed and Rerank are billed as separate per-token/per-call add-ons from generation — easy to under-forecast a RAG pipeline's total bill by pricing only the chat model.

What an audit finds on Cohere API

Teams budget using Command R7B's list price but route real RAG/agent traffic through Command A for quality, quietly increasing effective $/token by roughly 65x without updating cost forecasts.

Cohere API pricing FAQ

Is Cohere's cheapest model good enough for production?

Command R7B is about 65x cheaper than Command A per output token, but most teams needing strong reasoning or tool-use end up on Command A, which requires a sales conversation for full production rate limits.

Does Cohere charge for embeddings and rerank separately from chat?

Yes — Embed and Rerank are billed independently from the generation model, so a full RAG stack's real cost is generation + embed + rerank combined, not just the chat rate.

What is Cohere API really costing you?

Spendassay connects your Cohere API spend to actual usage and engineering output — a proof-level AI cost report with the wasted dollars named. Neutral across every AI vendor.

Free · read-only · no card

More: the complete AI cost management guide · every tool, with prices · cost per outcome · compare platforms · how we measure