How Much Does Kimi K3 Cost to Use?
Kimi K3 costs $0.30 per million cached input tokens, $3.00 per million uncached input tokens and $15.00 per million output tokens on Moonshot AI's own platform, against a context window of 1,048,576 tokens.
Those are Moonshot's own published list prices, and a sharp step up from the rest of the family. Output on Kimi K3 costs 3.75 times what the K2 line charges for the same million tokens.
| Model | Cached input / 1M | Uncached input / 1M | Output / 1M | Context |
|---|---|---|---|---|
| Kimi K3 | $0.30 | $3.00 | $15.00 | 1,048,576 |
| Kimi K2.7-Code | $0.19 | $0.95 | $4.00 | 262,144 |
| Kimi K2.7-Code highspeed | $0.38 | $1.90 | $8.00 | 262,144 |
| Kimi K2.6 | $0.16 | $0.95 | $4.00 | 262,144 |
Moonshot also prices K2.6 batch jobs at 60% of the standard rate, a useful lever if your bulk work fits that model.
Almost nobody pays a frontier bill out of revenue in year one. AI Perks tracks $7.7M in credits across 194 companies.

Three Kimi Models Shipped in 98 Days
Moonshot AI released Kimi K2.6, Kimi K2.7-Code and Kimi K3 inside a 98-day window between April 20 and July 27, 2026, and all three price cards are still published side by side.
The sequence matters if you are choosing what to build on:
- April 20, 2026 - Kimi K2.6. A 1 trillion parameter mixture-of-experts model with 32 billion active parameters, released open source under a Modified MIT license, with a 300-sub-agent "Agent Swarm" system for long-horizon coding.
- June 12, 2026 - Kimi K2.7-Code. A coding-specialised successor. Moonshot reports +21.8% on its own internal Kimi Code Bench v2 and 30% fewer reasoning tokens than K2.6. Both numbers come from Moonshot's own testing rather than an independent benchmark.
- July 16, 2026 - Kimi K3 in chat and through the API, with open weights following on July 27, 2026. At 2.8 trillion parameters it is the largest open-weight model released to date, has native vision, and entered the Artificial Analysis leaderboard at number 3.
The useful detail is that K3 did not replace anything. K2.6 and K2.7-Code are still listed at the rates in the table above, so the cheap tier is still on the menu.
Where Kimi K3 Actually Runs
Three independent hosts shipped Kimi K3 inference on July 27, 2026, the day the weights went public: Fireworks AI, Modal and Baseten.
A day-zero three-way race makes "where to run Kimi K3 cheapest" a real question. Everything published so far:
| Host | Cached input / 1M | Input / 1M | Output / 1M | Notes |
|---|---|---|---|---|
| Moonshot AI (first party) | $0.30 | $3.00 | $15.00 | 1,048,576-token context |
| Fireworks AI (Standard) | $0.30 | $3.00 | $15.00 | Fast tier is exactly 1.5x: $0.45 / $4.50 / $22.50. US-hosted adds 10% |
| Baseten | Not published | $3.00 listed | Unclear, confirm on Baseten's page | Day-zero inference and training API |
| Modal | Not published | Token-based, rate not published | Not published | Standing $30/month general compute credit |
Two things fall out of that table. Fireworks matches Moonshot's first-party rate to the cent on its Standard tier, so the choice between them is latency, region and throughput rather than price.
And two of the three hosts have not published a complete rate card. Modal confirmed token-based billing without naming a per-million figure, and Baseten's output price did not render cleanly on its own pricing table. Any comparison showing four tidy numbers is filling in blanks the vendors left empty.

Does Kimi K3 Have a Free Tier?
No. Moonshot AI publishes no permanent free tier for Kimi K3, Kimi K2.7-Code or Kimi K2.6.
The API is reported to stay silent until the account carries a balance, with a minimum around $1. That figure is reported rather than documented, so treat it as unconfirmed and fund the account before your first call either way.
Two adjacent free routes exist, and neither comes from Moonshot:
- Fireworks AI made Kimi K3 Fast with the full 1M-token context the featured model on its Fire Pass free tier, a way to try K3 without pre-funding anything.
- Modal publishes a standing $30 per month general compute credit, not a K3 launch promotion, so it applies to whatever you run there.
Neither funds a production workload on a 2.8 trillion parameter model. That gap between a trial and a real budget is what startup credit programs close, and AI Perks tracks which ones are open now.
Kimi K3 or Kimi K2.7-Code: What the Price Gap Buys
Kimi K3 costs roughly 3.3 times more per task than Kimi K2.7-Code, so K3 has to earn the difference with context length rather than quality alone.
Price one agentic job: 300,000 uncached input tokens in, 30,000 output tokens back.
| Model | Input cost | Output cost | Cost per task | At 1,000 tasks/month |
|---|---|---|---|---|
| Kimi K3 | $0.90 | $0.45 | $1.35 | $1,350 |
| Kimi K2.7-Code highspeed | $0.57 | $0.24 | $0.81 | $810 |
| Kimi K2.7-Code | $0.29 | $0.12 | $0.41 | $405 |
| Kimi K2.6 | $0.29 | $0.12 | $0.41 | $405 |
A spread of roughly $945 a month at that volume is what you weigh, and the honest tiebreaker is context. Kimi K3 holds 1,048,576 tokens against 262,144 on the K2.7 line, exactly four times as much. If your work fits inside 262,144 tokens, K2.7-Code is hard to argue against. If you feed a whole repository in one pass, K3 is the only model here that can take it.
Caching is the other lever, and it pays better on K3 than anywhere else in the range. A cache hit costs one tenth of a miss on K3 ($0.30 against $3.00), versus about one fifth on K2.7-Code. A stable system prompt and a consistent prefix turn that $1.35 task into roughly $0.54.

What the Kimi K3 License Actually Allows
Kimi K3's weights are open but the license is custom, not a standard open source license, and companies above US$20 million in revenue must negotiate a commercial contract with Moonshot before reselling access to the model.
For most readers this clause never bites. Calling the API, or serving the model inside your own product, is not reselling access. For anyone building an inference-reselling business on the weights, it is the first paragraph to read.
Do not carry the terms across generations either. Kimi K2.6 went out under a Modified MIT license, a materially different posture from the custom K3 license. Same vendor, same year, two legal regimes.
Self-hosting is the other thing people underestimate. A 2.8 trillion parameter mixture-of-experts model is not a weekend deployment, and three hosts launched day-zero APIs because renting capacity is the realistic option.
How to Cover a Kimi K3 Bill With Free AI Credits
No dedicated Moonshot AI startup credit program has surfaced for Kimi K3, so the working route is credits at the layer underneath and at the providers beside it.
| Credit category | What it pays for |
|---|---|
| Cloud and GPU compute | The machines an open-weight model runs on |
| Inference platforms and hosts | Serving Kimi K3 without owning GPUs |
| Frontier model APIs | Model calls at the providers beside Moonshot |
| Developer tooling | Everything wrapped around the model |
Total tracked: $7.7M in credits across 194 companies. Current terms for every category sit at AI Perks.
Two moves work when your model has no credit program of its own:
Fund the layer underneath. Kimi K3 runs on rented GPUs at Fireworks, Modal or Baseten, and Modal alone publishes a standing $30 per month compute credit. Compute credits reach the same destination from a different direction.
Hold model credits elsewhere. Free credits at other providers release cash for the one workload that genuinely needs a 1M-token window. Moonshot shipped three models in 98 days and credit programs turn over on a similar clock, so re-check quarterly.

Frequently Asked Questions
How much does Kimi K3 cost per million tokens?
Kimi K3 costs $0.30 per million cached input tokens, $3.00 per million uncached input tokens and $15.00 per million output tokens on Moonshot AI's platform, with a 1,048,576-token context window. Fireworks AI matches that exact rate on its Standard tier. Credits that offset bills like this are tracked at getaiperks.com.
Is Kimi K3 free to use?
No. Moonshot AI publishes no permanent free tier for Kimi K3. Fireworks AI features Kimi K3 Fast with 1M context on its Fire Pass free tier, and Modal publishes a standing $30 per month general compute credit. Neither offer comes from Moonshot, and neither replaces a funded budget for sustained use.
Where is the cheapest place to run Kimi K3?
Moonshot AI and Fireworks AI both publish $3.00 input and $15.00 output per million tokens, so on list price they tie. Modal and Baseten have not published complete rate cards yet. Decide on latency, region and throughput, then cut the bill with credits from getaiperks.com.
Is Kimi K3 open source?
The weights are open, but the license is custom rather than a recognised open source license. Companies above US$20 million in revenue must negotiate a commercial contract with Moonshot before reselling access to the model. Kimi K2.6 shipped separately under a Modified MIT license, so the two carry genuinely different terms.
Should I use Kimi K3 or Kimi K2.7-Code?
Use Kimi K2.7-Code when your work fits inside its 262,144-token context, because it costs roughly a third as much per task. Use Kimi K3 when you need the 1,048,576-token window, which is four times larger. Moonshot reports K2.7-Code at +21.8% on its own internal benchmark, a vendor-run figure rather than an independent one.
Can free AI credits pay for Kimi K3 usage?
Indirectly, yes. Moonshot AI runs no dedicated startup credit program, so the route is compute credits at whichever host serves your Kimi K3 traffic, plus model credits elsewhere that free up cash. AI Perks tracks $7.7M across 194 companies, including the inference platforms serving K3 today.
The largest open-weight model ever released still sends you an invoice. Arrange for someone else to pay it.