AI cost · Together AI

Together AI cost: what you actually pay

Together's 100+ model catalog is priced individually and fine-tuning rates vary more than 10x by base model, so one DeepSeek/Llama/Qwen tier's sticker rate says little about what a fine-tune or a different-parameter sibling model will cost.

Public pricing

Token rates (USD per 1M tokens)

ModelInputOutputCache read
DeepSeek V4 Proflagship reasoning model$1.74$3.48$0.2
Llama 3.3 70Bflat input=output rate; popular current Llama tier$1.04$1.04
Qwen3.5 9Bsmall/cheap tier$0.17$0.25
GPT-OSS-120B$0.15$0.6

Serverless per-token pricing coexists with dedicated GPU rentals ($3.99-$8.99/hr on-demand, cheaper on 1-6 month reservations) — at high sustained volume, dedicated GPU-hour pricing can undercut per-token serverless cost, but the crossover point isn't shown on the pricing page. Cached-input discounts (e.g., ~$0.20/M on DeepSeek V4 Pro) are only published for a subset of larger/MoE models, not the whole catalog.

Last updated: 2026-07-22. Sourced from Together AI Pricing (official).

Where the real Together AI cost hides

The sticker is the smallest part of the story. For Together AI, these are the line items that quietly inflate the bill:

  • Fine-tuning token pricing varies more than 10x by base model — a DeepSeek-R1-class or Kimi K2 LoRA fine-tune carries a $20-60/M-token minimum charge that can dwarf the inference cost it's meant to optimize.
  • Cached-input discounts are only listed for a subset of models (mostly large MoE/reasoning models) — smaller dense models like Llama 3.3 70B have no published cache discount.
  • Reserved GPU clusters bill by the hour regardless of utilization; committing to a 91-180 day reservation for spiky/unpredictable workloads locks in cost that per-token serverless pricing would have avoided.

What an audit finds on Together AI

High-volume teams default to serverless per-token pricing long after their steady-state usage would be cheaper on a reserved GPU-hour plan (or vice versa) — the crossover point isn't obvious from the pricing page and requires modeling actual tokens/sec.

Together AI pricing FAQ

Is Together AI cheaper than Fireworks for the same open model?

Very close — both price named models like DeepSeek V4 Pro ($1.74/$3.48) and GPT-OSS-120B ($0.15/$0.60) almost identically as of July 2026, so model choice matters more than vendor choice for these.

Does Together offer a batch-job discount?

Cached-input pricing (not a separate 'batch' toggle) is available on select large/MoE models, cutting input cost 80-90%; most smaller dense models don't have a published cache rate.

What is Together AI really costing you?

Spendassay connects your Together AI spend to actual usage and engineering output — a proof-level AI cost report with the wasted dollars named. Neutral across every AI vendor.

Free · read-only · no card

More: the complete AI cost management guide · every tool, with prices · cost per outcome · compare platforms · how we measure