AI cost · Together AI
Together AI cost: what you actually pay
Together's 100+ model catalog is priced individually and fine-tuning rates vary more than 10x by base model, so one DeepSeek/Llama/Qwen tier's sticker rate says little about what a fine-tune or a different-parameter sibling model will cost.
Public pricing
Token rates (USD per 1M tokens)
| Model | Input | Output | Cache read |
|---|---|---|---|
| DeepSeek V4 Proflagship reasoning model | $1.74 | $3.48 | $0.2 |
| Llama 3.3 70Bflat input=output rate; popular current Llama tier | $1.04 | $1.04 | — |
| Qwen3.5 9Bsmall/cheap tier | $0.17 | $0.25 | — |
| GPT-OSS-120B | $0.15 | $0.6 | — |
Serverless per-token pricing coexists with dedicated GPU rentals ($3.99-$8.99/hr on-demand, cheaper on 1-6 month reservations) — at high sustained volume, dedicated GPU-hour pricing can undercut per-token serverless cost, but the crossover point isn't shown on the pricing page. Cached-input discounts (e.g., ~$0.20/M on DeepSeek V4 Pro) are only published for a subset of larger/MoE models, not the whole catalog.
Last updated: 2026-07-22. Sourced from Together AI Pricing (official) ↗.
Where the real Together AI cost hides
The sticker is the smallest part of the story. For Together AI, these are the line items that quietly inflate the bill:
- Fine-tuning token pricing varies more than 10x by base model — a DeepSeek-R1-class or Kimi K2 LoRA fine-tune carries a $20-60/M-token minimum charge that can dwarf the inference cost it's meant to optimize.
- Cached-input discounts are only listed for a subset of models (mostly large MoE/reasoning models) — smaller dense models like Llama 3.3 70B have no published cache discount.
- Reserved GPU clusters bill by the hour regardless of utilization; committing to a 91-180 day reservation for spiky/unpredictable workloads locks in cost that per-token serverless pricing would have avoided.
What an audit finds on Together AI
High-volume teams default to serverless per-token pricing long after their steady-state usage would be cheaper on a reserved GPU-hour plan (or vice versa) — the crossover point isn't obvious from the pricing page and requires modeling actual tokens/sec.
Together AI pricing FAQ
Is Together AI cheaper than Fireworks for the same open model?
Very close — both price named models like DeepSeek V4 Pro ($1.74/$3.48) and GPT-OSS-120B ($0.15/$0.60) almost identically as of July 2026, so model choice matters more than vendor choice for these.
Does Together offer a batch-job discount?
Cached-input pricing (not a separate 'batch' toggle) is available on select large/MoE models, cutting input cost 80-90%; most smaller dense models don't have a published cache rate.
What is Together AI really costing you?
Spendassay connects your Together AI spend to actual usage and engineering output — a proof-level AI cost report with the wasted dollars named. Neutral across every AI vendor.
Free · read-only · no card
Other api platforms pricing
More: the complete AI cost management guide · every tool, with prices · cost per outcome · compare platforms · how we measure