AI cost · Fireworks

Fireworks AI cost: what you actually pay

The headline per-token rate only applies to a model's default serverless endpoint at standard priority — switching to 'Fast'/'Priority' tiers or an uncommon long-tail open model quietly moves you onto a pricier per-parameter-count tier.

Public pricing

Token rates (USD per 1M tokens)

ModelInputOutputCache read
DeepSeek V4 Proflagship reasoning model$1.74$3.48$0.145
Kimi K2.6$0.95$4$0.16
Qwen3.7 Plusrepresentative current Qwen tier$0.4$1.6$0.08
GPT-OSS-120Bcheap popular open model; long-tail/unnamed open models instead bill via parameter-count tiers ($0.10-$1.20/M)$0.15$0.6$0.015

Batch inference is 50% of serverless pricing; 'Fast'/'Priority' service-tier variants of the same model cost roughly 1.5-2x standard tier. Models without an individually negotiated rate fall back to parameter-count tiers ($0.10/M under 4B up to $1.20/M for 56-176B MoE), applied equally to input and output.

Last updated: 2026-07-22. Sourced from Fireworks AI Pricing (official), Fireworks Serverless Pricing Docs.

Where the real Fireworks AI cost hides

The sticker is the smallest part of the story. For Fireworks AI, these are the line items that quietly inflate the bill:

  • 'Fast' and 'Priority' variants of the same model can cost 1.5-2x the standard serverless rate — easy to select unintentionally when copying a model ID from docs or a benchmark.
  • Unnamed/long-tail open models fall back to parameter-count tiers rather than a cheaper named-model rate, so swapping to a slightly different checkpoint can silently change the bill.
  • Fine-tuning training-token pricing is separate and scales steeply by base model (up to $40/M tokens for full-parameter DPO on 300B+ models) — a fine-tune job can cost more than the inference it's meant to cheapen.

What an audit finds on Fireworks AI

Prompt caching and batch discounts (50% off each) require deliberate opt-in via specific API parameters/endpoints; teams calling the standard synchronous endpoint with repeated system prompts pay full price for tokens that could be discounted.

Fireworks AI pricing FAQ

Does Fireworks charge the same for every open model of a given size?

No — popular checkpoints (DeepSeek, Kimi, Qwen, GLM, GPT-OSS) get individually negotiated rates, often cheaper than the generic parameter-count tier applied to less common models.

Is there a cheaper way to run large batch workloads?

Yes, Fireworks' batch API is 50% of serverless pricing, but it is a separate endpoint/flow from the standard chat completions call, so it must be adopted deliberately.

What is Fireworks AI really costing you?

Spendassay connects your Fireworks AI spend to actual usage and engineering output — a proof-level AI cost report with the wasted dollars named. Neutral across every AI vendor.

Free · read-only · no card

More: the complete AI cost management guide · every tool, with prices · cost per outcome · compare platforms · how we measure