AI cost · Together AI
Together AI cost: what you actually pay
Together's 100+ model catalog is priced individually and fine-tuning rates vary more than 10x by base model, so one DeepSeek/Llama/Qwen tier's sticker rate says little about what a fine-tune or a different-parameter sibling model will cost.
Public pricing
Token rates (USD per 1M tokens)
| Model | Input | Output | Cache read |
|---|---|---|---|
| DeepSeek V4 Proflagship reasoning model | $1.74 | $3.48 | $0.2 |
| Llama 3.3 70Bflat input=output rate; popular current Llama tier | $1.04 | $1.04 | — |
| Qwen3.5 9Bsmall/cheap tier | $0.17 | $0.25 | — |
| GPT-OSS-120B | $0.15 | $0.6 | — |
Serverless per-token pricing coexists with dedicated GPU rentals ($3.99-$8.99/hr on-demand, cheaper on 1-6 month reservations) — at high sustained volume, dedicated GPU-hour pricing can undercut per-token serverless cost, but the crossover point isn't shown on the pricing page. Cached-input discounts (e.g., ~$0.20/M on DeepSeek V4 Pro) are only published for a subset of larger/MoE models, not the whole catalog.
Last updated: 2026-07-22. Sourced from Together AI Pricing (official) ↗.
Where the real Together AI cost hides
The sticker is the smallest part of the story. For Together AI, these are the line items that quietly inflate the bill:
- Fine-tuning token pricing varies more than 10x by base model — a DeepSeek-R1-class or Kimi K2 LoRA fine-tune carries a $20-60/M-token minimum charge that can dwarf the inference cost it's meant to optimize.
- Cached-input discounts are only listed for a subset of models (mostly large MoE/reasoning models) — smaller dense models like Llama 3.3 70B have no published cache discount.
- Reserved GPU clusters bill by the hour regardless of utilization; committing to a 91-180 day reservation for spiky/unpredictable workloads locks in cost that per-token serverless pricing would have avoided.
What an audit finds on Together AI
High-volume teams default to serverless per-token pricing long after their steady-state usage would be cheaper on a reserved GPU-hour plan (or vice versa) — the crossover point isn't obvious from the pricing page and requires modeling actual tokens/sec.
Worked example: Together AI at production volume
Take 50M input and 10M output tokens a month — a modest production workload, and a round number we picked — on DeepSeek V4 Pro at the published rates above.
- Monthly
- $122
- Annual
- $1,462
- On Qwen3.5 9B
- $11
Same volume, $1,330 a year apart. Most workloads are a mix — some of it genuinely needs the frontier model and some of it does not — so the real question is what share of your calls is the routine kind, and that is a measurement, not a guess.
Arithmetic on the published prices above. Your mix will differ — run your own numbers.
Together AI pricing FAQ
Is Together AI cheaper than Fireworks for the same open model?
Very close — both price named models like DeepSeek V4 Pro ($1.74/$3.48) and GPT-OSS-120B ($0.15/$0.60) almost identically as of July 2026, so model choice matters more than vendor choice for these.
Does Together offer a batch-job discount?
Cached-input pricing (not a separate 'batch' toggle) is available on select large/MoE models, cutting input cost 80-90%; most smaller dense models don't have a published cache rate.
How much can routing cut a Together AI bill?
On 50M input and 10M output tokens a month, DeepSeek V4 Pro costs $122 at published rates and Qwen3.5 9B costs $11 — $1,330 a year apart for the same volume. What you can actually move depends on how much of your traffic is routine, which is a measurement rather than a guess.
Are these Together AI prices current?
They were last checked against Together AI Pricing (official) on 2026-07-22. Every price on this page carries that date and a link to its source, and a scheduled check fails our build when any figure goes more than 30 days unverified. Vendors do change pricing between checks — the source link is there so you can confirm before you sign.
What is Together AI really costing you?
Spendassay connects your Together AI spend to actual usage and engineering output — a proof-level AI cost report with the wasted dollars named. Neutral across every AI vendor.
Free · read-only · no card
Other api platforms pricing
More: the complete AI cost management guide · every tool, with prices · cost per outcome · compare platforms · how we measure