AI cost · Fireworks
Fireworks AI cost: what you actually pay
The headline per-token rate only applies to a model's default serverless endpoint at standard priority — switching to 'Fast'/'Priority' tiers or an uncommon long-tail open model quietly moves you onto a pricier per-parameter-count tier.
Public pricing
Token rates (USD per 1M tokens)
| Model | Input | Output | Cache read |
|---|---|---|---|
| DeepSeek V4 Proflagship reasoning model | $1.74 | $3.48 | $0.145 |
| Kimi K2.6 | $0.95 | $4 | $0.16 |
| Qwen3.7 Plusrepresentative current Qwen tier | $0.4 | $1.6 | $0.08 |
| GPT-OSS-120Bcheap popular open model; long-tail/unnamed open models instead bill via parameter-count tiers ($0.10-$1.20/M) | $0.15 | $0.6 | $0.015 |
Batch inference is 50% of serverless pricing; 'Fast'/'Priority' service-tier variants of the same model cost roughly 1.5-2x standard tier. Models without an individually negotiated rate fall back to parameter-count tiers ($0.10/M under 4B up to $1.20/M for 56-176B MoE), applied equally to input and output.
Last updated: 2026-07-22. Sourced from Fireworks AI Pricing (official) ↗, Fireworks Serverless Pricing Docs ↗.
Where the real Fireworks AI cost hides
The sticker is the smallest part of the story. For Fireworks AI, these are the line items that quietly inflate the bill:
- 'Fast' and 'Priority' variants of the same model can cost 1.5-2x the standard serverless rate — easy to select unintentionally when copying a model ID from docs or a benchmark.
- Unnamed/long-tail open models fall back to parameter-count tiers rather than a cheaper named-model rate, so swapping to a slightly different checkpoint can silently change the bill.
- Fine-tuning training-token pricing is separate and scales steeply by base model (up to $40/M tokens for full-parameter DPO on 300B+ models) — a fine-tune job can cost more than the inference it's meant to cheapen.
What an audit finds on Fireworks AI
Prompt caching and batch discounts (50% off each) require deliberate opt-in via specific API parameters/endpoints; teams calling the standard synchronous endpoint with repeated system prompts pay full price for tokens that could be discounted.
Worked example: Fireworks AI at production volume
Take 50M input and 10M output tokens a month — a modest production workload, and a round number we picked — on DeepSeek V4 Pro at the published rates above.
- Monthly
- $122
- Annual
- $1,462
- On GPT-OSS-120B
- $14
Same volume, $1,300 a year apart. Most workloads are a mix — some of it genuinely needs the frontier model and some of it does not — so the real question is what share of your calls is the routine kind, and that is a measurement, not a guess.
Arithmetic on the published prices above. Your mix will differ — run your own numbers.
Fireworks AI pricing FAQ
Does Fireworks charge the same for every open model of a given size?
No — popular checkpoints (DeepSeek, Kimi, Qwen, GLM, GPT-OSS) get individually negotiated rates, often cheaper than the generic parameter-count tier applied to less common models.
Is there a cheaper way to run large batch workloads?
Yes, Fireworks' batch API is 50% of serverless pricing, but it is a separate endpoint/flow from the standard chat completions call, so it must be adopted deliberately.
How much can routing cut a Fireworks AI bill?
On 50M input and 10M output tokens a month, DeepSeek V4 Pro costs $122 at published rates and GPT-OSS-120B costs $14 — $1,300 a year apart for the same volume. What you can actually move depends on how much of your traffic is routine, which is a measurement rather than a guess.
Are these Fireworks AI prices current?
They were last checked against Fireworks AI Pricing (official) and Fireworks Serverless Pricing Docs on 2026-07-22. Every price on this page carries that date and a link to its source, and a scheduled check fails our build when any figure goes more than 30 days unverified. Vendors do change pricing between checks — the source link is there so you can confirm before you sign.
What is Fireworks AI really costing you?
Spendassay connects your Fireworks AI spend to actual usage and engineering output — a proof-level AI cost report with the wasted dollars named. Neutral across every AI vendor.
Free · read-only · no card
Other api platforms pricing
More: the complete AI cost management guide · every tool, with prices · cost per outcome · compare platforms · how we measure