Skip to main content
TokenCost logoTokenCost
Model ReleaseSeptember 10, 2026·13 min read

DeepSeek cut Flash to $0.15 this morning and will bill its Pro tier at the same rate from Monday. For the 96 hours in between, deepseek-v4-pro costs 4.4x more than the model DeepSeek says has already surpassed it.

DeepSeek V4.1 Flash went live at 04:00 UTC on September 10 at $0.15 input and $0.60 output off-peak, down from $0.22 and $0.66. The plan circulated the day before was to route every V4 Pro request to it in the same hour. A little over an hour after the new prices took effect, a second notice pushed that to September 14. The Pro card itself has not moved a cent, and there is no way to keep the Pro model after Monday.

Black and white railway tracks diverging at a switch under station lights at night

Photo by Valomukitse Arva-Zika on Unsplash

Flash input, off-peak

$0.22 → $0.15

31.8% off. Output moved 9.1%, cache hits 57.1%. Three lines, three different cuts.

Pro output, from Sep 14

$1.98 → $0.60

Not a price cut. The name deepseek-v4-pro will point at a different, smaller model and bill at its rate.

Three lines moved, none by the number you read

The new card is on DeepSeek's pricing page under a model called deepseek-flash, version DeepSeek-V4.1-Flash. The old card is on the same page as the Wayback Machine captured it at 16:47 UTC on September 9, under deepseek-v4-flash, version DeepSeek-V4-Flash-0731. We put them side by side so the percentages are ours and checkable, not somebody's rounding.

Per 1M tokensV4 Flash (Sep 9)V4.1 Flash (Sep 10)Change
Input, cache hit$0.007$0.003-57.1%
Input, cache miss$0.22$0.15-31.8%
Output$0.66$0.60-9.1%
Input, cache hit (peak)$0.014$0.006-57.1%
Input, cache miss (peak)$0.44$0.30-31.8%
Output (peak)$1.32$1.20-9.1%

So the headline cut is 31.8%, and it is the input line. Output barely moved: $0.66 to $0.60. The cache-hit line took the deepest cut, from $0.007 to $0.003, which restores the 98% cache discount that V4 Flash launched with in April and lost when the August 16 reprice narrowed it to 96.8%. If your traffic is mostly repeated system prompts and long shared prefixes, this is a bigger deal to you than the input line. If your traffic is mostly generation, the new card is 9% cheaper, and that is all.

You may have read 60%, 33% and 11% instead. Those figures are circulating in at least one English writeup and they are not wrong, exactly. They are the cuts on the yuan card.

Off-peak, per 1MV4 FlashV4.1 FlashCut in yuanCut in dollars
Input, cache hit¥0.05¥0.02-60.0%-57.1%
Input, cache miss¥1.5¥1-33.3%-31.8%
Output¥4.5¥4-11.1%-9.1%

Why do they differ? Because DeepSeek set the two currencies at two exchange rates. Divide the old Flash cache-miss or output line, or any current Pro line, yuan by dollars, and you get 6.82. Divide any new Flash line and you get 6.67. Both cards sit on the same page today. A customer paying in yuan got a 33.3% cut on input; a customer paying in dollars got 31.8%, because the new dollar lines were set 2.3% above their yuan equivalents at the rate every other line on the page still uses. It is a small thing. It is also the kind of small thing that means you should never take a percentage from a press summary when the two rate cards are a click away.

Pro is not being repriced. The name is being rewired.

Footnote (2) on the pricing page, in full: "After extensive testing, V4.1 Flash has comprehensively surpassed V4 Pro in performance, cost, speed, and total time, so we plan to retire V4 Pro in an orderly manner. From 12:00 Beijing Time on September 14, 2026, and until V4.1 Pro is released in the future, requests to deepseek-v4-pro will all be routed to V4.1 Flash and billed at the V4.1 Flash price."

Read that as a billing event and it is on a par with the 75% Pro cut of May, the one we wrote about and that lasted twelve weeks, except that this time the number behind the name changes too. Read it as a model event and it is a deprecation with a four-day notice period and no pin. Both readings are correct at once, which is what makes it hard to write about fairly.

deepseek-v4-pro, per 1MTo Sep 14 (V4 Pro)From Sep 14 (V4.1 Flash)ChangeRatio
Input, cache hit$0.022$0.003-86.4%7.33x
Input, cache miss$0.66$0.15-77.3%4.40x
Output$1.98$0.60-69.7%3.30x
Input, cache hit (peak)$0.044$0.006-86.4%7.33x
Input, cache miss (peak)$1.32$0.30-77.3%4.40x
Output (peak)$3.96$1.20-69.7%3.30x

The ratio column is the one to stare at. From August 16 until this morning the Pro premium over Flash was exactly 3.00x on input and on output. Now it is 4.40x on input, 3.30x on output and 7.33x on cache hits, and DeepSeek is charging that premium, until Monday, for a model it has stated in writing is worse on performance, cost, speed and total time than the cheaper one beside it. We cannot think of another vendor that has published that sentence about its own top tier while still billing for it.

Put a workload on it. Say 100M input tokens a month at a 60% cache-hit rate, plus 10M output tokens, all off-peak. On V4 Pro that is $1.32 for the hits, $26.40 for the misses and $19.80 for the output: $47.52. Route it to V4.1 Flash and the same three lines are $0.18, $6.00 and $6.00: $12.18, a 74.4% drop. The same workload on the old V4 Flash card was $15.82, so a Flash user saves 23.0% and a Pro user saves more than three times that, for doing nothing.

Two things the footnote does not say. Whether routed Pro traffic keeps Pro's concurrency limit of 500 or inherits Flash's 2,500; the rate limit page still lists both numbers against both names. And when V4.1 Pro arrives. The changelog says "the future release of V4.1 Pro" and nothing else, so the interval during which deepseek-v4-pro is a Flash model with a Pro name has a start and no end.

The date moved once already

If you read coverage published in the first hours of September 10 it says Pro routing began at 04:00 UTC that day, with under 48 hours of notice. That was the plan, and the notice period was real. It is no longer what is happening, and the reason it is not is a Hacker News thread.

UTCWhat happened
Sep 8, by 08:04Beta id deepseek-v4.1-flash-expires-on-0910 opens to DeepSeek's user group at V4 Flash prices, 20 concurrent requests per account.
Sep 9, by 11:19A group notice reposted to Hacker News: Flash prices change at 12:00 Beijing on Sep 10, and 'following the official launch of V4.1 Flash' all Pro requests route to Flash. The thread reaches 399 points.
Sep 10, 04:00V4.1 Flash goes live as deepseek-flash. New Flash prices take effect. deepseek-v4-flash and deepseek-v4-flash-vision-exp start routing to it.
Sep 10, by 05:11A second notice: V4 Pro discontinuation postponed to 12:00 Beijing on Sep 14. 'The price of V4 Pro remains unchanged during the service period.'
Sep 14, 04:00deepseek-v4-pro routes to V4.1 Flash and bills at the Flash price, 'until V4.1 Pro is released in the future'. No date for that.

None of the three notices was a changelog entry or a blog post. They were messages to DeepSeek's user group, reposted by members to Hacker News, where the September 9 one gathered 399 points and 209 comments in a day. The tone of the thread was not about the price. Simon Willison's comment was typical: "API model providers should lean towards not swapping out models on their paying customers, no matter how much 'better' the new model is meant to be." Others asked for a 404 instead of a silent reroute, or for a deprecation window like the ones Anthropic and OpenAI publish. Somewhere between that thread and 05:11 UTC the next morning, four days were added.

The postponement notice also carries a sentence worth reading twice: "If you continue to use our services after the billing adjustment, you will be deemed to have accepted the adjusted billing terms. If you do not agree, you may choose to cancel your service and apply for a refund." That is the whole opt-out. There is no snapshot id that keeps V4 Pro, no dated alias, no legacy endpoint. The only dated endpoint in DeepSeek's changelog was a V3.2 special that expired last December, and the beta id for this launch, deepseek-v4.1-flash-expires-on-0910, did what its name said on schedule.

One more thing about that date. 12:00 Beijing time is 04:00 UTC, and 04:00 UTC is the exact minute DeepSeek's first weekday peak window closes. The Pro tier ends at the end of a peak hour on a Monday morning. We doubt anyone chose that on purpose, but if you are pricing the last Pro requests you will ever make, the three hours before the switch cost double.

"Comprehensively surpassed" is fourteen of sixteen

DeepSeek's changelog lists nineteen scores for V4.1 Flash. Sixteen of them have a V4 Pro figure beside them in Table 3 of the V4.1 technical report; the three that do not are the vision benchmarks, which V4 Pro cannot take. Here are the sixteen, with the winner in the last column.

Benchmark (DeepSeek-reported)V4.1 FlashV4 Pro (0813)Higher
GPQA Diamond90.992.4Pro
HLE (no tools, text-only subset)39.142.7Pro
HLE (with tools)63.960.0Flash
Codeforces rating34713348Flash
MathArena Apex65.665.3Flash
Terminal-Bench 2.190.687.9Flash
Terminal-Bench 3.030.011.8Flash
Terminal-Bench 4.031.212.4Flash
DeepSWE v1.174.262.7Flash
ProgramBench20.315.5Flash
NL2Repo-Bench65.461.5Flash
CyberGym88.183.3Flash
SEC-Bench Pro62.856.4Flash
ExploitGym15.35.4Flash
Automation-Bench54.843.2Flash
Agents' Last Exam31.825.7Flash

Fourteen to two, and the fourteen include the agentic benchmarks that matter most for the kind of long tool-calling traffic that pays for a 1M-token context: Terminal-Bench 2.1 up 2.7 points, DeepSWE up 11.5, Automation-Bench up 11.6, and the two newer Terminal-Bench versions roughly two and a half times Pro's score. The two are GPQA Diamond, where Pro keeps a 1.5-point edge, and Humanity's Last Exam without tools, where Pro is 3.6 points ahead on the text-only subset that is the only like-for-like comparison the report offers. If your workload is hard scientific question answering with no tool access, the model you are being moved to on Monday is worse at it by DeepSeek's own numbers. That is a narrow case. It is also the case "comprehensively" was supposed to cover.

The reason the smaller model can do this is in the technical report, and it is a different design rather than a distillation. V4.1 Flash has 552B backbone parameters and, because of a causal encoder-decoder split, activates 16B per token when generating but only 8B when reading the prompt. V4 Flash was 284B with 13B active; V4 Pro is 1.6T with 49B active. The KV cache footprint is 890 bytes per token, about a quarter of V4 Flash's. So the new model is twice the size of the old Flash, a third the size of Pro, and cheaper to serve than either, which is how a price cut and a capability jump can land in the same release. The weights are on Hugging Face under MIT. For the first two hours after upload the README was titled DeepSeek-V4.1-Exp; a commit at 06:25 UTC renamed it, and nobody has said what the "Exp" was.

What nobody independent has measured yet is any of it. Artificial Analysis has no V4.1 Flash page as of this morning. Its current v4.3 index puts V4 Flash 0731 at 35 and V4 Pro 0813 at 36, one point apart, at $474.19 and $1,122.27 respectively to run the suite, so on the outgoing generation the independent view already said the Pro premium bought almost nothing. The speed figures going round, 355 to 427 tokens per second and 178 ms to first token, come from individual users' tests and appear in no DeepSeek document; the only measured baseline is Artificial Analysis's 65.7 tokens per second for V4 Pro and 122.6 for V4 Flash. We will update this post when a V4.1 Flash index score exists.

$0.15 is where three rate cards now meet

$0.15 per million input tokens is where DeepSeek landed, and it is also where Z.AI's GLM-5.3-Flash landed on September 9 when its launch discount lapsed, and where Alibaba's Qwen3.8-Flash has sat since late August. Three Chinese labs, one number, inside fifteen days. Nobody coordinated that. It is what a price floor looks like while it is forming.

ModelProviderInputCache hitOutput
DeepSeek V4.1 Flash (off-peak)DeepSeek$0.15$0.003$0.60
DeepSeek V4.1 Flash (peak)DeepSeek$0.30$0.006$1.20
GLM-5.3-FlashZ.AI$0.15$0.03$0.50
Qwen3.8-FlashAlibaba$0.15$0.015$0.47
GPT-5.6 LunaOpenAI$0.20$0.02$1.20
Gemini 3.5 Flash-LiteGoogle$0.30$0.03$2.50
Gemini 3.8 Flash (to Dec 31)Google$0.75$0.075$3.75
Claude Haiku 4.5Anthropic$1.00$0.10$5.00

On output DeepSeek is the dearest of the three at $0.60, against Qwen's $0.47 and GLM's $0.50. On cache hits it is not close: $0.003 is a fifth of Qwen's $0.015 and a tenth of GLM's $0.03, and neither of those two halves its card for seventeen hours a day. Which one is cheapest for you depends entirely on your hit rate and your shape. At a 90% cache-hit rate, off-peak, DeepSeek is cheaper once input outweighs output by more than about 12 to 1, which describes most agent loops; at 0% hits both Qwen and GLM are cheaper at every shape. Our calculator takes a cache-hit rate for exactly this reason.

The peak row is there because it is the row that gets left out. Between 01:00 and 04:00 and again between 06:00 and 10:00 UTC on a weekday, V4.1 Flash is $0.30 and $1.20, which puts it above GPT-5.6 Luna on input and level with it on output. A US West Coast team whose agents run overnight, or a European team that starts work at 07:00 UTC, spends much of its day in that row without ever seeing the $0.15. Google's Gemini 3.8 Flash is on this table at a promotional $0.75 that doubles on January 1; we covered that one last week. Fast, cheap and stable is still a pick-two.

Weekends went off-peak and nobody said so

While checking the old Flash card we read the peak-hours footnote in each Wayback capture, and it is not the same sentence throughout. On August 19 and August 22 it reads: "Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC (all other hours are off-peak)." From the August 24 capture onward it reads: "Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday (all other hours are off-peak)." The Chinese page carries the matching 周一至周五. The changelog for that week has an entry for the V4 Flash Vision Exp release on August 21 and nothing about pricing.

The arithmetic is not small. Seven peak hours a day, seven days a week, is 49 of 168 hours, or 29.2%; a flat workload paid 1.2917x the off-peak rate. Seven peak hours on five days is 35 of 168, or 20.8%, and the same workload pays 1.2083x. On the new Flash card that is a blended $0.18125 input and $0.725 output, against $0.19375 and $0.775 if weekends still counted. Every blended DeepSeek figure we published in August, in the peak/off-peak post and in our own pricing notes, used the 7-of-24 version, because that was the footnote on the day and nothing told us it had changed. Those figures have been about 7% too high since the fourth week of August.

We mention it not to flagellate but because it is the third silent edit to this one page since the summer: the peak schedule that was announced in June, published, and then removed in early August without ever being charged; the weekday clause added in late August; and now a Pro tier whose retirement lives in a footnote. DeepSeek's changelog is good when it is used. The pricing page changes more often than the changelog does.

What this did to our own catalogue

Four things. DeepSeek V4.1 Flash is now its own entry on the pricing page under the name DeepSeek uses for it, deepseek-flash, at $0.15 and $0.60 off-peak with the peak card and the cache price in its notes. We did not rename the old entry, because it carried Artificial Analysis scores measured on a model that no longer exists and the new one has no independent scores at all; V4 Flash and V4 Flash Vision Exp are marked retired instead, and their pages say where the names now route.

DeepSeek V4 Pro carries a dated price schedule: $0.66 and $1.98 through September 13, then $0.15 and $0.60, labelled for what it is, a routing to another model rather than a cut. Our calculator will flip on the day. Every DeepSeek blended figure now uses 35 of 168 hours. And Qwen3.8-Flash, which we had at $0.16 from Alibaba's launch blog, is corrected to the $0.15 on Alibaba's pricing page, which is the number a card is charged, and which is what puts three labs on one price instead of two.

If you have deepseek-v4-pro in production, the only decision available to you is whether to move to deepseek-flash yourself before Monday or let DeepSeek do it for you at 04:00 UTC. Moving yourself gets you the new price four days early and, more usefully, lets you run your evals against the model that will be answering next week while the old one is still there to compare against. After Monday there is nothing to compare against.