What Does Fireworks AI Actually Cost in 2026?
Fireworks AI serves Kimi K3 at $3.00 per million input tokens, $0.30 per million cached input tokens and $15.00 per million output tokens on its Standard serverless tier.
Those are Fireworks' published launch rates, and they match Moonshot AI's first-party price to the cent, which rarely happens with a third-party host.
Fireworks shipped Kimi K3 on July 27, 2026, the day Moonshot released the weights, inference and training together. Modal and Baseten shipped it the same day, making the launch a three-way race. Moonshot had opened K3 to chat and API users on July 16. It is a 2.8 trillion parameter mixture of experts, the largest open-weights release to date, and it debuted at number three on the Artificial Analysis leaderboard.
| Model and host | Input / 1M | Cached input / 1M | Output / 1M | Context |
|---|---|---|---|---|
| Kimi K3, Fireworks Standard | $3.00 | $0.30 | $15.00 | 1M |
| Kimi K3, Fireworks Fast | $4.50 | $0.45 | $22.50 | 1M |
| Kimi K3, Moonshot direct | $3.00 | $0.30 | $15.00 | 1,048,576 |
| Kimi K2.7-Code, Moonshot direct | $0.95 | $0.19 | $4.00 | 262,144 |
| Kimi K2.7-Code highspeed, Moonshot | $1.90 | $0.38 | $8.00 | 262,144 |
| Kimi K2.6, Moonshot direct | $0.95 | $0.16 | $4.00 | 262,144 |
Input is the cache-miss rate, cached input the cache-hit rate. Cached input costs 10% of uncached input on Kimi K3, so prompt caching is worth more here than almost any other optimisation. The generational jump is steep too: K3 output costs 3.75 times K2.7-Code output, for 2.8T parameters against K2.6's 1T with 32B active.
Nobody funds a frontier inference bill out of revenue in year one. AI Perks tracks which model and compute vendors are handing out credits.

Is Fire Pass a Real Free Tier?
Fire Pass is Fireworks' free tier, and its featured free model is now Kimi K3 Fast at the full 1M-token context. What it lacks is a published set of limits you can plan a product around.
A free tier fronted by a frontier model is unusual, and the contrast with the lab is sharp: Moonshot runs no permanent free tier on any Kimi model, and the account reportedly has to be pre-funded, with a minimum of about $1, before the API answers at all. Fireworks putting the same weights behind a free tier is the cheapest legitimate way to point a coding agent at a 1M-context model.
The caution is symmetric. Write-ups of Fire Pass circulate very specific conditions: invite codes, banned production workloads, revocation clauses, daily caps. None of it is confirmed in Fireworks' own launch material as checked here, so treat every restriction and number as unconfirmed until it appears in your own console.
That gap between a developer trial and a funded budget is what startup credit programmes close. AI Perks tracks which are open now.
The US Hosting Premium: Two Numbers That Disagree
Fireworks sells a US-hosted variant of Kimi K3 above the global price. The premium is reported two ways, and the gap between them is a factor of five.
Fireworks' own Kimi K3 launch material puts the US-hosted variant at 10% above the base serverless rate. A later summary of Fireworks' serverless pricing documentation reports a policy instead: from September 1, 2026, newly launched US-only models are priced at 1.5 times base, with one exception, GLM-5.2 Fast US, said to match global pricing. That second account is a secondary summary, not confirmed against Fireworks' documentation here.
Both can hold at once: one describes a model priced before that date, the other models launched after it. Read the US price off your own console before committing to a data residency clause, because on a frontier model 10% versus 50% is the difference between a rounding error and a line item.
The speed tier, by contrast, is not in dispute. On a single run of 1M input and 100K output tokens with no cache hits:
| Route | Input cost | Output cost | Run total |
|---|---|---|---|
| Kimi K3, Fireworks Standard | $3.00 | $1.50 | $4.50 |
| Kimi K3, Fireworks Fast | $4.50 | $2.25 | $6.75 |
Fast costs exactly 1.5 times Standard across all three token types, $2.25 more on that run: cheap for latency you need, expensive as a default.

Serverless Training: Reported, Not Confirmed
Fireworks is reported to have launched a Serverless Training API alongside Kimi K3, billing LoRA fine-tuning per token rather than by the reserved GPU hour. Almost everything about it is second-hand.
The design change is the interesting part, and it does not depend on a rate card. Reserved training charges for GPU hours whether you are learning anything or not. Per-token training charges for work done, which suits the LoRA-sized experiments most teams actually run. Reported coverage at launch spanned Qwen 3.5 9B, Qwen 3.6 27B and Kimi K3.
Everything below is reported by a secondary source and was not confirmed on Fireworks' own documentation in this pass. Treat it as a shape, not a price:
- Billing is said to split across prefill, cached prefill, sample and train stages rather than one flat per-token rate.
- The reported range across those stages is $0.66 to $32.55 per 1M tokens. A 49x spread makes any single headline number meaningless: a quote built on the wrong stage is wrong by an order of magnitude.
- Per-model LoRA rates, sample run costs and preview status circulate in launch coverage, none of them confirmed here.
Budget a fine-tuning programme from your own console output, not from launch coverage. Credits are the cheaper variable to solve for, and AI Perks tracks compute and model credits side by side.
Fireworks, Modal, Baseten or Moonshot Direct?
On published numbers, Fireworks Standard and Moonshot direct are a tie, and the other two day-zero hosts have not published enough to compare.
Fireworks' Standard tier and Moonshot's platform both list $3.00 input, $0.30 cached input and $15.00 output per million tokens at a 1M-token context. That parity removes price from the argument entirely.
The rest of the field is thinner than the launch coverage suggests:
- Baseten lists Kimi K3 input at $3.00 per 1M. Its published output figure is unclear enough to be worth reading off Baseten's own pricing page.
- Modal confirmed token-based pricing for Kimi K3 but published no per-1M rate in its own launch post. Modal does give new workspaces $30 a month of general compute credit, a standing offer rather than a Kimi launch perk.
- Moonshot direct runs no permanent free tier on any Kimi model, and prices K2.6 batch jobs at 60% of the standard rate, the one discount here that is first-party and unambiguous.
One caution on the weights: Kimi K3's custom licence requires companies above US$20M in revenue to negotiate a commercial contract with Moonshot before reselling access. That constrains a self-hosted copy, not what you call through somebody else's API.

What Actually Cuts the Fireworks Bill
Fire Pass covers a developer laptop. Credit programmes cover a product. They are different instruments and you want both.
Inference credits and compute credits reach the same destination from different directions. Fireworks bills for tokens; the GPU layer underneath runs its own startup programmes. Worth knowing first: no dedicated startup credit programme was found on Moonshot's own Kimi platform for K2.6, K2.7-Code or K3, so the credit layer for this family sits with the hosts, not the lab.
The levers that survive scrutiny, largest first:
- Cache first. At 10% of the uncached rate, prompt caching on Kimi K3 is a larger lever than switching providers, and the only one that costs nothing to try.
- Right-size the model. K2.7-Code at $0.95 in and $4.00 out is roughly a quarter of K3's output price, and Moonshot reports it scoring 21.8% higher than K2.6 on the lab's own Kimi Code Bench v2 while using 30% fewer reasoning tokens. Internal benchmarks are marketing, but fewer reasoning tokens show up on the invoice regardless.
- Batch what is not latency-sensitive. Moonshot prices K2.6 batch jobs at 60% of the standard rate.
- Re-check quarterly. Fireworks' US pricing policy reportedly changed on September 1, 2026 with no model launch attached. Rate cards move faster than articles about rate cards, including this one.
Frequently Asked Questions
Does Fireworks AI have a free tier?
Yes. Fire Pass is Fireworks' free tier, and its featured free model is Kimi K3 Fast at a 1M-token context. Caps, invite conditions and workload restrictions circulate widely but are not confirmed in the material checked here, so verify them in your own console. For budgets beyond personal development, credits are tracked at getaiperks.com.
How much does Kimi K3 cost on Fireworks AI?
Kimi K3 costs $3.00 per million input tokens, $0.30 per million cached input tokens and $15.00 per million output tokens on the Standard serverless tier. The Fast tier lists at $4.50 / $0.45 / $22.50, exactly 1.5 times Standard. A US-hosted variant carries a premium whose size is reported inconsistently.
Is Fireworks AI cheaper than using Kimi directly?
No, and not more expensive either. Fireworks' Standard tier matches Moonshot AI's published rate to the cent on all three token types, both at a 1M-token context. Choose on region, latency and model breadth, then cut whichever bill you end up with using credits from getaiperks.com.
What is the Fireworks US hosting premium?
It is reported two ways. Fireworks' Kimi K3 launch material puts the US-hosted variant 10% above the base serverless price. A secondary summary of Fireworks' pricing documentation instead reports 1.5 times base for US-only models launched from September 1, 2026, with GLM-5.2 Fast US as an exception at global pricing. Check the model you actually plan to call before budgeting either figure.
Is the Fireworks Serverless Training API generally available?
Fireworks has not published a complete public rate card for serverless training, and figures circulating from secondary sources, including a reported $0.66 to $32.55 per 1M token range across billing stages, are unconfirmed. Price any fine-tuning programme from your own console and offset it with credits from getaiperks.com.
Are there startup credits for Kimi K3?
No dedicated startup credit programme was found on Moonshot's side for K2.6, K2.7-Code or K3. On the host side, Modal offers new workspaces $30 a month of general compute credit, which is not Kimi-specific. Current terms across the inference and compute layers are tracked at getaiperks.com.
Run the frontier model. Let someone else pay for the tokens.