DeepSeek V4-Pro Pricing: What It Costs vs V4.1-Flash

DeepSeek-V4-Pro-0813 pricing per million tokens, the peak and off-peak split, the cache-hit discount, and why model choice beats scheduling.

DeepSeekDeepSeek V4-ProLLM API PricingOff-Peak PricingAI Perks
Author Avatar
Andrew
AI Perks Team
10,145

Quick Answer

DeepSeek-V4-Pro-0813 costs $0.66 per million cache-miss input tokens off-peak and $1.32 at peak, with output at $1.98 off-peak and $3.96 at peak. A cache hit drops input to $0.022 off-peak. DeepSeek publishes no API free tier and runs no startup credit program, unlike the 194 companies tracked at getaiperks.com.

What Does DeepSeek-V4-Pro Cost Per Million Tokens?

DeepSeek-V4-Pro-0813 costs $0.66 per million cache-miss input tokens off-peak and $1.32 at peak. Output is $1.98 off-peak and $3.96 at peak. A cache hit drops input to $0.022 off-peak.

DeepSeek-V4-Pro-0813 reached general availability on 13 August 2026 as the flagship of the V4 line, built for tool use, code execution and multi-step autonomous workflows. It is priced by the clock: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, and every other hour costs half.

Rates published in DeepSeek's API documentation, in USD per million tokens:

ModelToken typeOff-peakPeak
DeepSeek-V4-Pro-0813Input, cache hit$0.022$0.044
DeepSeek-V4-Pro-0813Input, cache miss$0.66$1.32
DeepSeek-V4-Pro-0813Output$1.98$3.96
DeepSeek-V4.1-FlashInput, cache hit$0.003$0.006
DeepSeek-V4.1-FlashInput, cache miss$0.15$0.30
DeepSeek-V4.1-FlashOutput$0.60$1.20

At the April 2026 preview the V4 pair was reported to carry a 1M-token context window with up to 384K tokens of output, a switchable thinking mode, and open weights under an MIT licence. That comes from launch coverage, not DeepSeek's current pricing page, so confirm it against the exact model string you call. AI Perks tracks pricing and credit programs across 194 companies, and DeepSeek stands out on the second count: it publishes no startup credit program.


Round Funded
SponsoredRaise money from 10,000+ active vetted investors.
Start Raising

Peak Is Only 21% of the Week

Most coverage frames off-peak as a discount you unlock. The arithmetic says the opposite: peak covers 7 hours a weekday, 35 out of 168, so 79% of the week already bills at the lower rate.

Off-peak is the default. Peak is a surcharge you pay by accident, in a window falling across the European morning. You are not hunting a discount, you are avoiding a 35 hour penalty box, and the fix is to check that no cron job landed there. Interactive traffic is the exception, because that clock is set by your users.


Why V4.1-Flash Beats V4-Pro on Price

Moving a job off-peak halves the bill, exactly 2x. Moving it from V4-Pro to V4.1-Flash cuts cache-miss input by 4.4x and output by 3.3x. Model choice is the bigger lever, and the two compound.

One million cache-miss input tokens plus one million output tokens, at both ends of the clock:

Model and timingInputOutputTotal
V4-Pro at peak$1.32$3.96$5.28
V4-Pro off-peak$0.66$1.98$2.64
V4.1-Flash at peak$0.30$1.20$1.50
V4.1-Flash off-peak$0.15$0.60$0.75

V4-Pro in its cheapest hour still costs 1.76x what V4.1-Flash costs in its most expensive hour. No scheduling trick closes that gap.

Stack both levers and the spread widens again: V4-Pro at peak against V4.1-Flash off-peak is 7x on identical token counts, and only one of those two decisions is set by the clock. The usual advice to push batch work into cheap hours assumes the model is fixed. It rarely is, and demoting it pays even when traffic cannot be rescheduled.

Flash wins on price in every hour of the day, peak included. V4-Pro's advantage is capability alone, so it has to earn its place on task difficulty.


Round Funded
SponsoredRaise money from 10,000+ active vetted investors.
Start Raising

Why a V4-Pro Agent Run Costs More Than the Table Suggests

Output costs 3x what cache-miss input costs on V4-Pro, and the workloads DeepSeek built V4-Pro for - agentic loops, tool calls, multi-step execution - generate the most output per unit of work.

A month of real usage: 100M input tokens at a 70% cache hit rate plus 20M output tokens, at the rates above:

ConfigurationInputOutputMonthly total
V4-Pro, peak$42.68$79.20$121.88
V4-Pro, off-peak$21.34$39.60$60.94
V4.1-Flash, peak$9.42$24.00$33.42
V4.1-Flash, off-peak$4.71$12.00$16.71

Three things follow.

Output is 65% of the V4-Pro off-peak bill. Once caching works and jobs avoid peak, DeepSeek stops being an input problem and becomes a response-length one.

The full spread is 7.3x. Same tokens, same vendor, same month. Top row to bottom row is two decisions.

The cache discount is steeper on Flash. A V4-Pro cache hit is 30x cheaper than a miss; on V4.1-Flash it is 50x. Prompt-prefix discipline pays off more on the cheap model, the opposite of what teams assume.


What DeepSeek Has and Has Not Published

DeepSeek publishes per-token pricing and little else a buyer wants. There is no stated API free tier on the official pricing page, no startup credit program, and no separately published pricing for the April preview models.

Any guide written before mid-August describes a pricing model that no longer exists:

DateWhat shippedPricing status
24 April 2026V4-Pro and V4-Flash previewNot separately published at preview
31 July 2026DeepSeek-V4-Flash-0731, stableReported at $0.14 input / $0.28 output, flat
13 August 2026DeepSeek-V4-Pro-0813, general availabilitySuperseded by the split below
Mid-August 2026Peak and off-peak pricing introducedCurrent structure
Around 10 September 2026DeepSeek-V4.1-FlashCurrent Flash rates

The April preview details, the July Flash figures and the exact date of the change are reported, not confirmed on DeepSeek's current pricing page: treat them as history, not budget inputs. The V4-Pro and V4.1-Flash rates above come from DeepSeek directly.

The August change, as reported, moved Flash from a flat $0.14 cache-miss input to the current $0.15 off-peak and $0.30 at peak. That is not a price cut with an asterisk but a price split: run off-peak and you pay roughly the old rate, run at peak and you pay about double.

On free access: the DeepSeek chat app is free. Claims that new API accounts receive free credits are inconsistently reported and not confirmed on the official pricing page, so treat any figure you see elsewhere as unverified. DeepSeek runs no startup credit program either, leaving no free runway at all. AI Perks tracks $7.7M in credits across 194 companies that do run one.

Find the providers that actually give credits at getaiperks.com →


Round Funded
SponsoredRaise money from 10,000+ active vetted investors.
Start Raising

Where the Off-Peak Discount Does Not Apply

Time-of-day pricing is a first-party DeepSeek mechanic. Third-party hosts and routers set their own rates, and a flat rate that looks cheap on paper can sit at DeepSeek's peak price around the clock.

Baseten is reported to have added DeepSeek V4.1 Flash to its Model APIs on 10 September 2026, with input listed at $0.30 per million tokens. Neither that figure nor the matching output figure is confirmed here (the output number was not reliably readable when checked), so verify both with the vendor before budgeting. Taken at face value, $0.30 is exactly DeepSeek's peak input rate and double its off-peak rate, charged every hour of the day. Baseten's standing new-workspace trial credit is reported at $30, which is neither DeepSeek-specific nor new.

The trade can still be correct: you buy latency, region and simplicity, and a flat rate forecasts more easily than a schedule. Price it honestly though, because a rate set at the peak number never gets the 79% of the week that would otherwise bill at half. Free runway comes from providers that fund it, and that list is at getaiperks.com.


How to Cut a DeepSeek-V4-Pro Bill, in Order of Impact

Lever 1: Fund the expensive tiers elsewhere. DeepSeek gives no credits, so use getaiperks.com to cover the premium providers you route hard tasks to, and let DeepSeek carry volume work.

Lever 2: Demote every task that does not need the flagship. On the workload above, V4-Pro to V4.1-Flash is 3.6x cheaper at identical settings. Nothing else here comes close.

Lever 3: Audit your schedule for accidental peak usage. Peak is 35 hours a week, so check that no cron job landed in 01:00 to 04:00 or 06:00 to 10:00 UTC on a weekday.

Lever 4: Stabilise your prompt prefix. Cache hits need an identical leading prefix, so put volatile content last. That is where the 30x on V4-Pro and the 50x on Flash live.

Lever 5: Attack response length. After the four above, output is roughly two thirds of the bill and the largest thing left.


Round Funded
SponsoredRaise money from 10,000+ active vetted investors.
Start Raising

Frequently Asked Questions

How much does DeepSeek-V4-Pro cost per million tokens?

Cache-miss input is $0.66 off-peak and $1.32 at peak. Output is $1.98 off-peak and $3.96 at peak. A cache hit costs $0.022 off-peak and $0.044 at peak. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.

What are DeepSeek's peak hours and how much of the week are they?

Peak is 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. That is 7 hours a weekday, 35 out of the 168 in a week, or about 21%. The remaining 79%, including all weekend, bills at half the peak rate.

Is DeepSeek-V4-Pro worth it over V4.1-Flash?

Only for work that genuinely needs the flagship. V4-Pro costs 4.4x more on cache-miss input, 3.3x more on output, and about 3.6x more on a mixed monthly workload. Route by task difficulty first, then optimise the schedule.

Does DeepSeek give free credits for V4-Pro?

No. The DeepSeek chat app is free, but the official pricing page states no API free tier and there is no startup credit program. Reports of free credits for new API accounts are unconfirmed. For real free runway, AI Perks tracks $7.7M across 194 companies.

Do third-party hosts pass on DeepSeek off-peak pricing?

Generally no. Time-of-day pricing is a first-party mechanic, and hosts set their own rates. Baseten is reported to list DeepSeek V4.1 Flash input at $0.30 per million, which would equal DeepSeek's peak rate charged all day, though that is not confirmed here. Verify the current rate with the host before you commit.

What happened to the deepseek-v4-flash endpoint name?

DeepSeek-V4.1-Flash retired the deepseek-v4-flash and deepseek-v4-flash-vision-exp model names, and requests now route to the current Flash model at Flash pricing. Code written against the old identifiers needs updating, so pin the current model string rather than relying on aliases.


Subscribe at getaiperks.com →

DeepSeek prices low and funds nothing. The providers that do fund you are at getaiperks.com.

This content is for informational purposes only and may contain inaccuracies. Credit programs, amounts, and eligibility requirements change frequently. Always verify details directly with the provider.