What Does DeepSeek-V4-Pro Cost Per Million Tokens?
DeepSeek-V4-Pro-0813 costs $0.66 per million cache-miss input tokens off-peak and $1.32 at peak. Output is $1.98 off-peak and $3.96 at peak. A cache hit drops input to $0.022 off-peak.
DeepSeek-V4-Pro-0813 reached general availability on 13 August 2026 as the flagship of the V4 line, built for tool use, code execution and multi-step autonomous workflows. It is priced by the clock: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, and every other hour costs half.
Rates published in DeepSeek's API documentation, in USD per million tokens:
| Model | Token type | Off-peak | Peak |
|---|---|---|---|
| DeepSeek-V4-Pro-0813 | Input, cache hit | $0.022 | $0.044 |
| DeepSeek-V4-Pro-0813 | Input, cache miss | $0.66 | $1.32 |
| DeepSeek-V4-Pro-0813 | Output | $1.98 | $3.96 |
| DeepSeek-V4.1-Flash | Input, cache hit | $0.003 | $0.006 |
| DeepSeek-V4.1-Flash | Input, cache miss | $0.15 | $0.30 |
| DeepSeek-V4.1-Flash | Output | $0.60 | $1.20 |
At the April 2026 preview the V4 pair was reported to carry a 1M-token context window with up to 384K tokens of output, a switchable thinking mode, and open weights under an MIT licence. That comes from launch coverage, not DeepSeek's current pricing page, so confirm it against the exact model string you call. AI Perks tracks pricing and credit programs across 194 companies, and DeepSeek stands out on the second count: it publishes no startup credit program.

Peak Is Only 21% of the Week
Most coverage frames off-peak as a discount you unlock. The arithmetic says the opposite: peak covers 7 hours a weekday, 35 out of 168, so 79% of the week already bills at the lower rate.
Off-peak is the default. Peak is a surcharge you pay by accident, in a window falling across the European morning. You are not hunting a discount, you are avoiding a 35 hour penalty box, and the fix is to check that no cron job landed there. Interactive traffic is the exception, because that clock is set by your users.
Why V4.1-Flash Beats V4-Pro on Price
Moving a job off-peak halves the bill, exactly 2x. Moving it from V4-Pro to V4.1-Flash cuts cache-miss input by 4.4x and output by 3.3x. Model choice is the bigger lever, and the two compound.
One million cache-miss input tokens plus one million output tokens, at both ends of the clock:
| Model and timing | Input | Output | Total |
|---|---|---|---|
| V4-Pro at peak | $1.32 | $3.96 | $5.28 |
| V4-Pro off-peak | $0.66 | $1.98 | $2.64 |
| V4.1-Flash at peak | $0.30 | $1.20 | $1.50 |
| V4.1-Flash off-peak | $0.15 | $0.60 | $0.75 |
V4-Pro in its cheapest hour still costs 1.76x what V4.1-Flash costs in its most expensive hour. No scheduling trick closes that gap.
Stack both levers and the spread widens again: V4-Pro at peak against V4.1-Flash off-peak is 7x on identical token counts, and only one of those two decisions is set by the clock. The usual advice to push batch work into cheap hours assumes the model is fixed. It rarely is, and demoting it pays even when traffic cannot be rescheduled.
Flash wins on price in every hour of the day, peak included. V4-Pro's advantage is capability alone, so it has to earn its place on task difficulty.

Why a V4-Pro Agent Run Costs More Than the Table Suggests
Output costs 3x what cache-miss input costs on V4-Pro, and the workloads DeepSeek built V4-Pro for - agentic loops, tool calls, multi-step execution - generate the most output per unit of work.
A month of real usage: 100M input tokens at a 70% cache hit rate plus 20M output tokens, at the rates above:
| Configuration | Input | Output | Monthly total |
|---|---|---|---|
| V4-Pro, peak | $42.68 | $79.20 | $121.88 |
| V4-Pro, off-peak | $21.34 | $39.60 | $60.94 |
| V4.1-Flash, peak | $9.42 | $24.00 | $33.42 |
| V4.1-Flash, off-peak | $4.71 | $12.00 | $16.71 |
Three things follow.
Output is 65% of the V4-Pro off-peak bill. Once caching works and jobs avoid peak, DeepSeek stops being an input problem and becomes a response-length one.
The full spread is 7.3x. Same tokens, same vendor, same month. Top row to bottom row is two decisions.
The cache discount is steeper on Flash. A V4-Pro cache hit is 30x cheaper than a miss; on V4.1-Flash it is 50x. Prompt-prefix discipline pays off more on the cheap model, the opposite of what teams assume.
What DeepSeek Has and Has Not Published
DeepSeek publishes per-token pricing and little else a buyer wants. There is no stated API free tier on the official pricing page, no startup credit program, and no separately published pricing for the April preview models.
Any guide written before mid-August describes a pricing model that no longer exists:
| Date | What shipped | Pricing status |
|---|---|---|
| 24 April 2026 | V4-Pro and V4-Flash preview | Not separately published at preview |
| 31 July 2026 | DeepSeek-V4-Flash-0731, stable | Reported at $0.14 input / $0.28 output, flat |
| 13 August 2026 | DeepSeek-V4-Pro-0813, general availability | Superseded by the split below |
| Mid-August 2026 | Peak and off-peak pricing introduced | Current structure |
| Around 10 September 2026 | DeepSeek-V4.1-Flash | Current Flash rates |
The April preview details, the July Flash figures and the exact date of the change are reported, not confirmed on DeepSeek's current pricing page: treat them as history, not budget inputs. The V4-Pro and V4.1-Flash rates above come from DeepSeek directly.
The August change, as reported, moved Flash from a flat $0.14 cache-miss input to the current $0.15 off-peak and $0.30 at peak. That is not a price cut with an asterisk but a price split: run off-peak and you pay roughly the old rate, run at peak and you pay about double.
On free access: the DeepSeek chat app is free. Claims that new API accounts receive free credits are inconsistently reported and not confirmed on the official pricing page, so treat any figure you see elsewhere as unverified. DeepSeek runs no startup credit program either, leaving no free runway at all. AI Perks tracks $7.7M in credits across 194 companies that do run one.
Find the providers that actually give credits at getaiperks.com →

Where the Off-Peak Discount Does Not Apply
Time-of-day pricing is a first-party DeepSeek mechanic. Third-party hosts and routers set their own rates, and a flat rate that looks cheap on paper can sit at DeepSeek's peak price around the clock.
Baseten is reported to have added DeepSeek V4.1 Flash to its Model APIs on 10 September 2026, with input listed at $0.30 per million tokens. Neither that figure nor the matching output figure is confirmed here (the output number was not reliably readable when checked), so verify both with the vendor before budgeting. Taken at face value, $0.30 is exactly DeepSeek's peak input rate and double its off-peak rate, charged every hour of the day. Baseten's standing new-workspace trial credit is reported at $30, which is neither DeepSeek-specific nor new.
The trade can still be correct: you buy latency, region and simplicity, and a flat rate forecasts more easily than a schedule. Price it honestly though, because a rate set at the peak number never gets the 79% of the week that would otherwise bill at half. Free runway comes from providers that fund it, and that list is at getaiperks.com.
How to Cut a DeepSeek-V4-Pro Bill, in Order of Impact
Lever 1: Fund the expensive tiers elsewhere. DeepSeek gives no credits, so use getaiperks.com to cover the premium providers you route hard tasks to, and let DeepSeek carry volume work.
Lever 2: Demote every task that does not need the flagship. On the workload above, V4-Pro to V4.1-Flash is 3.6x cheaper at identical settings. Nothing else here comes close.
Lever 3: Audit your schedule for accidental peak usage. Peak is 35 hours a week, so check that no cron job landed in 01:00 to 04:00 or 06:00 to 10:00 UTC on a weekday.
Lever 4: Stabilise your prompt prefix. Cache hits need an identical leading prefix, so put volatile content last. That is where the 30x on V4-Pro and the 50x on Flash live.
Lever 5: Attack response length. After the four above, output is roughly two thirds of the bill and the largest thing left.

Frequently Asked Questions
How much does DeepSeek-V4-Pro cost per million tokens?
Cache-miss input is $0.66 off-peak and $1.32 at peak. Output is $1.98 off-peak and $3.96 at peak. A cache hit costs $0.022 off-peak and $0.044 at peak. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.
What are DeepSeek's peak hours and how much of the week are they?
Peak is 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. That is 7 hours a weekday, 35 out of the 168 in a week, or about 21%. The remaining 79%, including all weekend, bills at half the peak rate.
Is DeepSeek-V4-Pro worth it over V4.1-Flash?
Only for work that genuinely needs the flagship. V4-Pro costs 4.4x more on cache-miss input, 3.3x more on output, and about 3.6x more on a mixed monthly workload. Route by task difficulty first, then optimise the schedule.
Does DeepSeek give free credits for V4-Pro?
No. The DeepSeek chat app is free, but the official pricing page states no API free tier and there is no startup credit program. Reports of free credits for new API accounts are unconfirmed. For real free runway, AI Perks tracks $7.7M across 194 companies.
Do third-party hosts pass on DeepSeek off-peak pricing?
Generally no. Time-of-day pricing is a first-party mechanic, and hosts set their own rates. Baseten is reported to list DeepSeek V4.1 Flash input at $0.30 per million, which would equal DeepSeek's peak rate charged all day, though that is not confirmed here. Verify the current rate with the host before you commit.
What happened to the deepseek-v4-flash endpoint name?
DeepSeek-V4.1-Flash retired the deepseek-v4-flash and deepseek-v4-flash-vision-exp model names, and requests now route to the current Flash model at Flash pricing. Code written against the old identifiers needs updating, so pin the current model string rather than relying on aliases.
DeepSeek prices low and funds nothing. The providers that do fund you are at getaiperks.com.