What Does a New Baseten Workspace Start With?
New Baseten workspaces are reported to start with a $30 trial credit, and it reads as a standing offer rather than a limited launch promotion.
That sentence carries a hedge most coverage drops. The $30 figure comes from reporting on Baseten's pricing and changelog surfaces rather than a headline offer restated in one obvious place, and the pricing page did not render consistently across repeat checks. So: reported, not vendor-confirmed. Check the balance your workspace actually shows.
What the credit applies to is clearer than the number. It sits on Baseten Model APIs, the shared endpoint layer hosting open-weight models such as Kimi K3, GLM-5.3 and DeepSeek V4.1 Flash, the same surface a production workload would call. The credit buys a real test, not a sandbox. AI Perks tracks the current terms alongside $7.7M in credits across 194 companies.

What Baseten Shipped Between July and September 2026
Baseten spent the quarter doing two things at once: adding frontier open-weight models to its Model APIs, and building products for the labs that make those models.
| Date | What launched | Who it is for |
|---|---|---|
| Jul 27, 2026 | Kimi K3 day-zero API | Developers |
| Jul 29, 2026 | Baseten for Model Labs | Model labs |
| Aug 28, 2026 | GLM-5.3 on Model APIs | Developers |
| Aug 29, 2026 | Frontier Gateway | Model labs |
| Sep 3, 2026 | GLM-5.3-Flash on fast capacity | Developers |
| Sep 10, 2026 | DeepSeek V4.1 Flash | Developers |
| Sep 13, 2026 | Six model deprecations announced | Everyone |
The Kimi K3 launch is the one worth understanding. Moonshot AI released an open-weight mixture-of-experts model, reported at 2.8 trillion to 3 trillion parameters with native vision and a 1M-token context window, and Fireworks AI, Modal and Baseten each shipped a hosted API the same day. Fireworks and Baseten shipped training APIs alongside inference. A three-way day-zero race means the same weights come from three vendors on three billing models, so provider choice is a pricing question, not a capability one.
Baseten Model APIs Pricing: Only Half the Table Reads Clean
Baseten publishes an input price per million tokens for each Model API. The output side is not something this page can quote with confidence.
| Model | Input per 1M tokens | Output per 1M tokens | Notes |
|---|---|---|---|
| Kimi K3 | $3.00 | Not reliably readable | Open-weight, native vision, 1M context |
| GLM-5.3 | $1.40 | Not reliably readable | 1M-token context |
| DeepSeek V4.1 Flash | $0.30 | Not reliably readable | 552B parameters, multimodal |
| GLM-5.3-Flash | Not confirmed | Not confirmed | Dedicated fast-serving capacity |
Those input rates are what Baseten's pricing table listed on checking. The output column did not render consistently across fetches, which is why no Baseten output rate appears above.
For scale on the missing half, look at the one provider publishing everything. Fireworks AI lists Kimi K3 at $3.00 input, $0.30 cached input and $15.00 output per million tokens on its Standard tier. Output costs five times input there. Any Baseten cost model built on the $3.00 input rate alone is wrong by an unknown multiple, and on generation-heavy workloads that multiple is most of the invoice. Treat any article printing a Baseten output price without a caveat as a guess. AI Perks tracks provider pricing as it moves.

Six Models Leave Baseten Model APIs on September 25, 2026
Baseten announced on September 13, 2026 that GLM-4.7, Kimi K2.7, Kimi K2.6, Inkling, Inkling-Small and DeepSeek V4 Pro retire from Model APIs on September 25, 2026.
This is the part that quietly breaks things. Three kinds of reader are affected.
Anyone with a model slug hardcoded in production. An OpenAI-compatible endpoint fails the same way whether the model was renamed or removed. Pin a replacement before the 25th, not after it.
Anyone maintaining a comparison post. Every listicle pricing DeepSeek V4 Pro or the Kimi K2.x line on Baseten goes wrong on September 25, and most will never be updated. Assume a pre-September roundup has dead rows.
Anyone evaluating on old benchmarks. If your harness targets Kimi K2.7 or K2.6, the successor in the catalog is Kimi K3 at a different price point, so last month's comparison does not carry forward.
Treat these pairings as editorial judgment rather than a vendor-published migration path: GLM-4.7 to GLM-5.3, the Kimi K2.x line to Kimi K3, DeepSeek V4 Pro to DeepSeek V4.1 Flash. Prices move in both directions across those jumps, which is exactly when a trial credit earns its keep. See what is currently available at AI Perks.
Where to Run Kimi K3: Baseten, Fireworks or Modal
Three providers shipped Kimi K3 on launch day, and only one publishes a complete per-token price you can plan against.
| Provider | Kimi K3 input per 1M | Full token pricing published | Free entry point |
|---|---|---|---|
| Fireworks AI | $3.00 Standard, $4.50 Fast | Yes | Fire Pass free tier, featuring Kimi K3 Fast |
| Baseten | $3.00 | No, output not readable | Trial credit, reported at $30, one time |
| Modal | Not published | No | $30 per month general compute for new workspaces |
Fireworks publishes the whole picture: $3.00 input, $0.30 cached and $15.00 output per million on Standard, the Fast tier at $4.50, $0.45 and $22.50, and a US-hosted variant at roughly a 10% premium. Its free tier, Fire Pass, switched to Kimi K3 Fast as its featured free model with the full 1M-token context, the cheapest way to learn whether the model suits your task. Baseten's input price matches at $3.00. Modal confirmed token-based pricing in its launch post without publishing a per-million rate.
The two $30 figures are a coincidence worth reading carefully: they are not the same offer. Modal's is a recurring monthly compute allowance covering compute generally rather than one model API. Baseten's is a one-time trial credit on a workspace. Across a three-month evaluation that is $90 against $30. Both are worth holding: AI Perks exists because credits at several layers beat a large one at a single vendor.

Baseten for Model Labs and Frontier Gateway Explained
These two 2026 launches are the ones most developers will find irrelevant, and knowing why saves an evaluation week.
Baseten for Model Labs, announced July 29, 2026, is a product line for closed-weight model labs that want to distribute and monetize their models through Baseten's infrastructure. It inverts the usual relationship: instead of Baseten hosting open-weight models for developers, the lab is the customer and the lab's users are the traffic.
Frontier Gateway, announced August 29, 2026, is a managed, multi-tenant API gateway that lets a lab stand up a production-grade inference API on Baseten's stack rather than build a serving layer itself.
Neither has published pricing, both are lab-facing rather than developer-facing, and a trial credit does not apply to either. Reporting on both is thin and medium-confidence, so treat any detailed feature list or latency benchmark found elsewhere as unconfirmed until the vendor states it. For a developer calling models the consequence is indirect: if labs adopt this route, more closed-weight models become reachable through endpoints resembling the ones you already call. Worth watching at AI Perks.
How to Make a Starter Credit Last
Confirm output pricing before extrapolating anything. Input rates alone understate real cost on generation-heavy workloads, and Baseten's output column is the one number this page cannot hand you. A budget built on $3.00 per million for Kimi K3 is a budget built on half the invoice.
Debug on the Flash-tier models. GLM-5.3-Flash and the DeepSeek Flash line exist for latency and cost rather than peak quality. Prompt plumbing, retry logic and schema bugs do not need a frontier model to surface, and swapping the slug afterwards is a one-line change.
Pin your model versions. Six models retire on September 25, 2026. Dated slugs stop a deprecation becoming an outage, and a trial credit is most useful while you re-benchmark replacements.
Hold credits at more than one layer. Inference, general compute and frontier model access bill separately. Baseten covers the first, Modal or a hyperscaler the second, a frontier lab the third.

Frequently Asked Questions
How much are free Baseten credits worth?
New workspaces are reported to start with a $30 trial credit, described as a standing offer rather than a promotion. The figure comes from reporting rather than a clearly stated vendor headline, so treat $30 as reported and confirm the balance your workspace shows. Current terms are tracked at getaiperks.com.
What does Baseten charge per token?
Input rates read clean: Kimi K3 at $3.00 per million tokens, GLM-5.3 at $1.40 and DeepSeek V4.1 Flash at $0.30. Output rates are not reliably readable from the public pricing table, and this page will not guess at them. Confirm output pricing in the console before modelling a monthly bill.
Which models is Baseten deprecating in September 2026?
GLM-4.7, Kimi K2.7, Kimi K2.6, Inkling, Inkling-Small and DeepSeek V4 Pro retire from Baseten Model APIs on September 25, 2026, announced on September 13. The nearest replacements in the catalog are GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash, though that pairing is editorial rather than a vendor-published migration path.
Is Baseten cheaper than Fireworks for Kimi K3?
On input tokens both list $3.00 per million, so they are level on the only number that compares cleanly. Fireworks also publishes $0.30 cached input and $15.00 output per million on its Standard tier. Until Baseten's output rate is confirmed, a full cost comparison is not possible in either direction.
What is Kimi K3 and why did three providers launch it at once?
Kimi K3 is Moonshot AI's open-weight mixture-of-experts model, reported at 2.8 trillion to 3 trillion parameters with native vision and a 1M-token context window, released July 27, 2026. Fireworks AI, Modal and Baseten each shipped a hosted API the same day. Open weights make that possible: no provider needs permission, so launch day becomes a price race instead of an exclusive.
Can I combine Baseten credits with other AI credits?
Yes, and it is the normal approach. Inference, general compute and frontier model credits are three separate bills with three separate vendors, and holding one does not block another. Baseten covers the inference layer, Modal or a hyperscaler covers compute, and a frontier lab covers direct model calls. AI Perks tracks $7.7M across 194 companies at getaiperks.com.
Pin your model versions. Let someone else pay for the tokens.