DeepSeek says a significant price increase is coming and will not say how significant. We measured the room it has: 2.14x before anything cheaper exists, 3.20x if you cache.
A three-sentence notice went up on DeepSeek's pricing page on August 6, as footnote number two under the rate table. no percentage, no per-model breakdown, no effective date. Every write-up since has led with the word significant because that is the only quantity in the notice, and significant is not a number you can put in a budget. So we did the arithmetic the other way round. Take a fixed job of 10 million input and 2 million output tokens. DeepSeek V4 Flash bills $1.96 for it today. The cheapest mainstream alternative, Qwen3.7 Flash, bills $4.198, and that is only true below 32K input. Which means DeepSeek can put its entire rate card up 2.14x and still be the cheapest serious model you can rent. Turn caching on at an 80% hit rate and the headroom stretches to 3.20x, because a 98% cache discount is worth more the more you use it. There is one number that fits inside that gap and it is DeepSeek's own: when its V3 promotion ended in February 2025, output went up 4x and cache hits 5x. Apply those same multipliers to today's card and V4 Flash lands at $5.04, past both alternatives, for the first time since it launched. Also here: why the peak-hour rows DeepSeek published in June have quietly vanished from both language versions of the page, why the CNY and USD cards imply two different exchange rates, and why this is the rare price rise you can walk away from without changing models.

Photo by Gaurav Yadav on Unsplash
The short version, before the tables
Nothing has changed on your invoice yet. The rates below are what DeepSeek charges this morning, and the notice is a warning rather than a change. What we can tell you that the notice cannot is how far it could go before switching away makes sense, and the honest answer is: further than most people assume, and further still if your prefixes cache well.
What DeepSeek told you
- Overall pricing is going up
- The increase will be significant
- Plan your usage accordingly
What it left out
- Any percentage, on any leg
- Which of the two models moves
- A date, even an approximate one
- A reason
Read the notice, it is only three sentences
Most of the coverage quotes the first clause and paraphrases the rest, which loses the part that matters. Here is the whole thing as it appears on the English pricing page:
"We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice."
The Chinese page carries the same three sentences: 计划近期整体上调 DeepSeek API 服务的定价,预计涨幅较大,请合理安排您的使用。具体方案以正式通知为准。 The word doing the work is 较大, which is closer to considerable than to enormous. It is a hedge either way.
Three things are absent and all three are the things you would need. There is no percentage. There is no statement about which of the three legs moves, and DeepSeek has form for moving them by wildly different amounts. And there is no date, only 近期.
Worth noting where it is not. DeepSeek keeps a dated news index for releases and pricing changes, and the newest entry there is still the V4 preview from April 24. This announcement exists as an undated footnote on a docs page. That is a deliberate choice about how loudly to say something, and it is why you will see the date given as both August 5 and August 6 in different places. August 6 is right: it is the Thursday that Chinese and English coverage both point at, and the August 5 stories are about a different DeepSeek headline entirely, which we will get to.
The one thing the notice does not need to spell out is what happens if you disagree with the new rates. DeepSeek's terms of service, effective April 29, 2026, already say fees may be adjusted, that you will be told through in-site notices, website announcements or email, and that continuing to use the service afterwards signifies acceptance. Refunds are lump-sum on unspent balance only.
What you are paying this morning
The catalogue is two models. The V3-era aliases, deepseek-chat and deepseek-reasoner, stopped answering after July 24, so if you are still pointing at those you have a bigger problem than a price rise. We covered that migration in the deprecation walkthrough.
| Model | Input, miss | Input, hit | Output | Context |
|---|---|---|---|---|
| deepseek-v4-flash | $0.14 | $0.0028 | $0.28 | 1M, 384K out |
| deepseek-v4-pro | $0.435 | $0.003625 | $0.87 | 1M, 384K out |
A small thing that tells you these are two separate price lists rather than one converted twice. The Chinese card reads 1 yuan input and 2 yuan output for Flash, 3 and 6 for Pro. Against the dollar figures that is an implied rate of 7.14 for Flash and 6.90 for Pro. Whichever currency your account bills in, you are on a card that was set by hand, and a card set by hand can be reset by hand.
The cache-hit rate is the number to keep an eye on. At $0.0028 against $0.14, DeepSeek discounts a repeated token by 98%. OpenAI and Anthropic take 90% off. Ling-3.0-flash and Qwen3.7 Flash take 80%. Nobody else in the cheap tier is close, and it is the single line most exposed to a rewrite, because it is the one that looks least like the rest of the market.
How far it can go before you would leave
Here is the calculation the notice makes impossible and also unnecessary. Instead of guessing the increase, price a fixed job on every plausible alternative, then read off the multiple. The job: 10 million input tokens, 2 million output, no caching, which is roughly the shape of a document-processing or extraction workload.
| Model | Input / 1M | Output / 1M | The job | vs DeepSeek |
|---|---|---|---|---|
| Ling-3.0-flash, promo card | $0.021 | $0.063 | $0.34 | 0.17x |
| V4 Flash, third-party host | $0.09 | $0.18 | $1.26 | 0.64x |
| deepseek-v4-flash, first party | $0.14 | $0.28 | $1.96 | 1.00x |
| Qwen3.7 Flash, under 32K in | $0.225 | $0.974 | $4.20 | 2.14x |
| gpt-5.6-luna | $0.20 | $1.20 | $4.40 | 2.24x |
| MiniMax-M3, under 512K in | $0.30 | $1.20 | $5.40 | 2.76x |
| deepseek-v4-pro | $0.435 | $0.87 | $6.09 | 3.11x |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $8.00 | 4.08x |
| Claude Haiku 4.5 | $1.00 | $5.00 | $20.00 | 10.20x |
| GLM-5.2 | $1.40 | $4.40 | $22.80 | 11.63x |
| Gemini 3.6 Flash | $1.50 | $7.50 | $30.00 | 15.31x |
| Kimi K3 | $3.00 | $15.00 | $60.00 | 30.61x |
Read the gap between rows three and four. DeepSeek could double its card and the invoice would still be smaller than Qwen's. That is not a comfortable position for anyone competing with it, and it is probably why the notice is worded the way it is: there is no commercial reason to be precise about a rise your customers have no realistic answer to.
Two rows are cheaper than DeepSeek already and both deserve an asterisk. Ling-3.0-flash at $0.021 is a promotional rate a router is funding rather than the host, with no published end date, and it scores 37.82 on Artificial Analysis' index against DeepSeek V4 Flash's 51.77. We took that card apart yesterday and would not budget on it. The third-party V4 Flash row is the interesting one and gets its own section below.
Everything from Haiku 4.5 down the table is a capability decision wearing a price tag. If you are on DeepSeek for cost, none of those rows is where you land, and pretending a 10x move is a lateral one helps nobody.
Cache well and the gap gets wider, not narrower
Run the same job with an 80% cache hit rate on the input leg, which is a normal figure for an agent replaying a stable prefix, and the ordering holds but the distances change. DeepSeek's 98% discount compounds against everyone else's 80% or 90%.
| Model | Cache discount | Blended input | The job | vs DeepSeek |
|---|---|---|---|---|
| deepseek-v4-flash | 98% | $0.03024 | $0.86 | 1.00x |
| deepseek-v4-pro | 99% | $0.08990 | $2.64 | 3.06x |
| Qwen3.7 Flash, under 32K in | 80% | $0.08100 | $2.76 | 3.20x |
| gpt-5.6-luna | 90% | $0.05600 | $2.96 | 3.43x |
| MiniMax-M3, under 512K in | 80% | $0.10800 | $3.48 | 4.04x |
| Claude Haiku 4.5 | 90% on reads | $0.28000 | $12.80 | 14.84x |
| GLM-5.2 | 81% | $0.48800 | $13.68 | 15.86x |
| Kimi K3 | 90% | $0.84000 | $38.40 | 44.53x |
Qwen3.7 Flash and gpt-5.6-luna swap places once caching is on, because a 90% discount on a $0.20 sticker beats an 80% discount on $0.225. Small mechanic, real money. And note what happened to DeepSeek's own Pro model: at 3.06x it is now cheaper than Qwen3.7 Flash, which it was not on the uncached job. If the increase lands unevenly across the two models, the sensible move for some workloads may be sideways within DeepSeek rather than out of it.
One honest caveat on this table. Gemini is missing from it, and deliberately: Google charges explicit cache storage by the hour on top of the token rate, so there is no single blended input figure to put in a column. Anthropic charges to write into the cache as well as read from it, so the Haiku row is a floor rather than a like-for-like figure. Cache pricing has three incompatible shapes in this market and any table that pretends otherwise is flattering somebody.
Last time a DeepSeek promotion ended, output went up 4x
The useful thing about this company is that it has done this before, twice, and both times it published the new numbers. That gives significant a range rather than a vibe.
| When | Direction | What moved |
|---|---|---|
| Feb 9, 2025 | Up | V3 promo ended. Output 4x, input miss 2x, cache hit 5x |
| Sep 5, 2025 | Up | New card at 16:00 UTC, off-peak discounts killed the same hour |
| Sep 29, 2025 | Down | V3.2-Exp, prices cut 50%+, effective immediately |
| May 23, 2026 | Down | V4-Pro 75% cut made permanent, a week before it was due to expire |
| Jun 30, 2026 | Up | Peak-hour doubling announced. Never appeared on an invoice |
| Aug 6, 2026 | Up | Significant, unspecified, undated |
The February 2025 row is the one to model, because it is the only time DeepSeek has published a new card materially above the old one, rather than letting a discount lapse back to a rate that was already on the books. Apply those exact multipliers to today's Flash card and you get $0.28 input, $0.014 cache hit, $1.12 output.
Price the job on that and it is $5.04 uncached, which is 2.57x today's bill and finally past both Qwen3.7 Flash at $4.20 and gpt-5.6-luna at $4.40. On the cached job it is $2.91 against luna's $2.96, a difference of under five cents on a twelve million token workload, which is a rounding error rather than a reason to stay.
So the 2025 precedent sits almost precisely on the line. A repeat of it is the smallest increase that would actually cost DeepSeek the price crown, and the largest one it could plausibly call significant without an exodus. If you want a single number to plan against, plan against 2.5x, and know that the cached leg is where it will hurt most because the 5x cache-hit multiplier was the steepest thing in that 2025 change and the 98% discount is the most out-of-line thing in this one.
The peak-hour rows are gone from the page
Here is a detail nobody covering this week has picked up, and we only caught it because we had quoted the page ourselves six days ago. On June 30 DeepSeek announced time-differentiated pricing: rates doubling during 09:00 to 12:00 and 14:00 to 18:00 Beijing time, the first time it had priced by clock rather than by token. We wrote it up at the time. On August 2, covering the 0731 build, we quoted the pricing page directly: it said DeepSeek would "soon adopt a peak/off-peak pricing policy" at double the regular rates, with the effective date still to be announced.
We checked both the English and Chinese pricing pages this morning. That sentence is not there. Two flat rows per model, no clock, no multiplier, and the Chinese page has no 高峰时段 section at all. A surcharge was announced, sat on the page for five weeks, never reached an invoice, and has now been removed without a word. In its place, a footnote saying the whole card is going up.
That is worth sitting with for a second, because it reframes the August notice. DeepSeek tried to raise prices six weeks ago using a mechanism that would have been invisible to anyone not watching the clock, decided against it, and has come back with a plain announcement that it is putting the whole card up instead. The quiet version did not happen. The loud version is the one they went with, and companies generally do not choose the loud version unless the number is large enough that the quiet version would not have covered it.
Why now, and what DeepSeek has not said
DeepSeek gave no reason. Not one sentence. Every compute-costs explanation you have read this week is a journalist's inference, and while the inference is reasonable it is not sourced, so we are going to label it rather than repeat it. The only rationale the company has ever offered on pricing was for the June peak-hour plan, where it talked about better distribution of resources, and that plan is the one it just deleted.
What we do have is a timing coincidence that is hard to ignore. On August 5, the day before the notice, DeepSeek V4 Flash was reported as the top model on OpenRouter's weekly ranking for July 27 to August 2, with 7.22 trillion tokens. The same article carries a second figure from a different platform: on OpenCode, V4 Flash handled 8 trillion tokens on August 1 alone, 5 trillion of it free-trial traffic and 3 trillion paid. Both numbers are single-sourced and we could not check either directly, because OpenRouter's rankings table renders client-side and returns nothing to a fetcher. Do not read the 5-in-8 split as DeepSeek-wide: it describes one integration's traffic, and free to the developer on OpenCode does not automatically mean unbilled to DeepSeek. It is still the most concrete pressure anyone has put a number on, and more concrete than any macro story about chips.
There is also a remark attributed to Liang Wenfeng in a transcript of investor comments published by Tencent Tech in July, which DeepSeek has not confirmed and which we are flagging as unverified for that reason: that he could raise the price by half, or double it, and token consumption would barely change. If the transcript is real, the notice reads less like a company under strain and more like one testing an elasticity it already believes it has. The table further up this page suggests he would be right.
You can leave without switching models
This is the part that makes this price rise different from every other one we have covered this year. When Anthropic put Sonnet 5 up 50% for September, the only way out was a different model. Here the model is MIT licensed and ungated on Hugging Face, so DeepSeek does not control who serves it.
Third-party hosts already sell V4 Flash below DeepSeek's own price. OpenRouter's models API returns $0.09 input, $0.18 output and $0.018 cache reads for deepseek-v4-flash-0731 at the full 1,048,576-token context, against DeepSeek's $0.14 and $0.28. That is 36% off the input leg on identical weights. The catch is in the third figure: $0.018 against $0.09 is an 80% cache discount, not 98%, which is exactly why that row slides from 0.64x to 0.79x once caching is on. Cheaper on paper, less cheap on an agent loop. And reseller cards carry discount badges and move without notice, so check it the day you need it rather than trusting this paragraph.
And if the increase is steep enough, renting the weights yourself becomes arithmetic rather than ideology. The GPU-hour maths in our Kimi K3 self-hosting breakdown transfers directly. The break-even there was brutal at DeepSeek's current rates and gets less brutal with every multiple.
What we would actually do this week
Not migrate. That is the main thing. A notice with no number and no date is not an event, and moving a working pipeline off the cheapest capable model on the market in anticipation of an increase you cannot size is how you spend three weeks to make your bill go up.
Do three cheaper things instead. Log your actual cache hit rate, because every number on this page moves with it and most teams have never looked. Rewrite whatever forecast currently assumes $0.14 to assume $0.35, which is 2.5x and lands where the 2025 precedent lands. And put a second provider behind a flag now, while nothing is urgent, so that the day the real card appears you are making a decision rather than an emergency.
One genuine risk to plan for that is not about money. DeepSeek prepays: if the new rates land while you have a balance sitting there, that balance buys fewer tokens than you budgeted, and the refund policy covers unspent lump sums rather than the difference. Keep the float small until the number is published.
We will update this post the day DeepSeek puts real numbers up. Until then you can run your own mix against every card in the tables above in the cost calculator, or line the alternatives up side by side on the pricing page.
Every number above, and where it came from
Given that half this post is about the difference between what DeepSeek said and what people have inferred, it would be poor form not to show our own working. Two claims here are single-sourced and one is from a leaked transcript. They are labelled as such below.
- DeepSeek: API pricing, English - The price-increase notice quoted in full, and the current card: deepseek-v4-flash at $0.14 cache miss, $0.0028 cache hit and $0.28 output, deepseek-v4-pro at $0.435, $0.003625 and $0.87, both 1M context and 384K max output. Fetched August 8, 2026, at which point the page carried no peak-hour or off-peak rows
- DeepSeek: API pricing, Chinese - The original wording of the notice, and the CNY card at 0.02, 1 and 2 yuan for Flash and 0.025, 3 and 6 for Pro, which is what the two different implied exchange rates are calculated from. Also confirms the absence of any 高峰时段 section
- DeepSeek: V3.1 release notes, August 2025 - The one previous increase DeepSeek dated precisely in its own words, September 5, 2025 at 16:00 UTC, with the off-peak discounts ending at the same moment. The news index it sits in is also where the August 2026 notice is not
- TechNode, August 6, 2026 - Dates the notice to Thursday August 6 and makes the point that the change is planned rather than in effect, with no schedule published
- TechNode, August 5, 2026 - The 7.22 trillion token weekly figure for July 27 to August 2 on OpenRouter, and separately the 8 trillion tokens on August 1 split 5 trillion free-trial and 3 trillion paid, which is an OpenCode figure rather than an OpenRouter one. Both single-sourced: OpenRouter's own rankings page renders client-side and could not be read directly, and TechNode itself notes the figures measure tracked platforms rather than DeepSeek's total usage
- The Next Web, August 6, 2026 - The clearest account of the trajectory from the V4-Pro cut through the peak surcharge to this notice. Its attribution of the increase to demand straining compute is the reporter's inference and is not sourced to DeepSeek
- DeepSeek: open platform terms of service - Section 6.2 on fee adjustment and notification, and that continued use after an adjustment signifies acceptance. Section 6.3 limits refunds to unspent balance in a lump sum. Effective April 29, 2026
- TechNode, June 30, 2026 - The peak-hour plan as announced: doubling during 09:00 to 12:00 and 14:00 to 18:00 Beijing time, described as DeepSeek's first time-differentiated API pricing. Those rows are no longer on the pricing page
- OpenAI: API pricing - gpt-5.6-luna at $0.20 input, $0.02 cached and $1.20 output with a 1.05M window. Note it tiers above 272K input tokens, where the whole request bills at $0.40 and $1.80, so the row above assumes you stay under that
- Alibaba Model Studio: Qwen3.7 Flash - $0.225 input, $0.045 cached and $0.974 output, valid only up to 32K input. Between 32K and 256K it is $0.749 and $2.998, above that $1.499 and $5.995, so the headroom figure is DeepSeek's worst case
- Anthropic: pricing - Claude Haiku 4.5 at $1 input, $5 output and $0.10 cache reads, with cache writes billed at 1.25x for five minutes or 2x for an hour, which is why the Haiku row is a floor
- Google: Gemini API pricing - Gemini 3.6 Flash at $1.50 and $7.50, Gemini 3.5 Flash-Lite at $0.30 and $2.50. Both charge cache storage per hour on top of the token rate, which is why neither appears in the blended cache table
- MiniMax: pay-as-you-go pricing - MiniMax-M3 at $0.30 input, $0.06 cached and $1.20 output up to 512K input, doubling above that. The listed rate is described as a permanent 50% discount on a $0.60 and $2.40 list
- OpenRouter: models API - deepseek/deepseek-v4-flash-0731 listed at $0.09 input, $0.18 output and $0.018 cache reads at 1,048,576 context, which is where the third-party figures and the 80% cache discount come from. Read from the API rather than the rendered page, and reseller rates move without notice
- Hugging Face: DeepSeek-V4-Flash-0731 - MIT licence, ungated weights and the serving documentation behind the claim that you can leave DeepSeek without leaving the model
- Liang Wenfeng investor transcript, via Tencent Tech, July 2026 - The remark about raising prices by half or doubling them without token consumption changing. Leaked rather than published, and DeepSeek has not confirmed the record, which is why it is labelled unverified above