Skip to main content
TokenCost logoTokenCost
IndustryAugust 16, 2026·10 min read

DeepSeek said the increase would be significant. It is 1.80x off-peak and 3.59x at peak, and only one of those two numbers fits inside the room we measured eight days ago.

The notice went up on August 6 as a three-sentence footnote with no percentage, no per-model breakdown and no date. Two of those three gaps have now been filled, and the percentage is still missing: every multiplier below is our arithmetic rather than DeepSeek's. At 16:00 UTC today DeepSeek's rate card splits in two, and which half you pay depends on what time your process happens to run. V4-Flash goes from a flat $0.14 and $0.28 per million tokens to $0.22 and $0.66 for seventeen hours of the day and $0.44 and $1.32 for the other seven. V4-Pro goes from $0.435 and $0.87 to $0.66 and $1.98 off-peak, $1.32 and $3.96 at peak. Off-peak is defined as exactly half of peak on every line, which makes this a single price rise wearing two hats rather than a discount scheme bolted onto an old card. The interesting part is not the headline multiplier. It is that the three lines on each card moved by three different amounts, the smallest increase landed on the line most people quote, and the largest landed on the one an agent loop actually lives on.

A large illuminated station clock face silhouetted against a city skyline at dusk

Photo by Fernando Mola-Davis on Unsplash

The billing day, hour by hour

Every hour of a UTC day below. The filled blocks are the seven hours DeepSeek now bills at double, and they are the only thing separating a job that costs $3.52 from an identical job that costs $7.04.

00:00 UTC06:0012:0018:0024:00
Off-peak, 17 hours:
$0.22 input and $0.66 output per million on V4-Flash. The reference job bills $3.52, which is 1.80x what it billed yesterday and still under everything comparable.
Peak, 7 hours:
$0.44 and $1.32. The same job bills $7.04, which is 3.59x yesterday and 60% more than gpt-5.6-luna charges at any hour.

The card DeepSeek published, in full

DeepSeek did not write a blog post about this. The announcement is a short changelog paragraph dated August 13 and footnote one under the rate table on the Models and Pricing page, and between them they carry the effective time, the peak windows, the halving rule and a six-cell table. Reproduced below with the outgoing rates alongside, because the footnote does not show you what you were paying yesterday and the multipliers are the whole story.

Per 1M tokensUntil todayOff-peakPeakOff-peak vs oldPeak vs old
V4-Flash input, cache miss$0.14$0.22$0.441.57x3.14x
V4-Flash input, cache hit$0.0028$0.007$0.0142.50x5.00x
V4-Flash output$0.28$0.66$1.322.36x4.71x
V4-Pro input, cache miss$0.435$0.66$1.321.52x3.03x
V4-Pro input, cache hit$0.003625$0.022$0.0446.07x12.14x
V4-Pro output$0.87$1.98$3.962.28x4.55x

Six rates became twelve, and not one of the six multipliers matches another. That is worth sitting with, because a price rise that multiplies every line by the same figure is a decision about revenue, and a price rise that moves six lines by six different amounts is a decision about which customers to keep. The input line everyone quotes in comparison tables moved least. The cache line almost nobody quotes moved most.

Cache is where the money actually moved

DeepSeek's cache discount has been the quiet reason its bills looked impossible. Most of the industry charges 10% of the input rate for a cache read. DeepSeek charged 2% on V4-Flash and 0.83% on V4-Pro, which is a 98% and a 99.17% discount respectively. Those two numbers are what let a long-context agent replay an enormous prompt every turn and barely register on the invoice.

Both are now 96.8% and 96.67%. Still well ahead of the field, but the gap to everyone else has narrowed, and on Pro it has collapsed: a cache read there went from 12x cheaper than the industry's standard 10% of input to about 3x cheaper. On Flash the move is milder, from 5x to 3.1x. The way it happened is worth naming precisely: DeepSeek did not raise cache reads in proportion to input. It raised input 1.57x and cache reads 2.50x on Flash, and input 1.52x and cache reads 6.07x on Pro. A V4-Pro cache read at peak costs 12.14 times what it cost yesterday, which is the single largest move anywhere on either card and roughly four times the increase on the input line sitting directly above it.

The practical consequence is that this increase lands hardest on exactly the workloads that had optimised hardest. Take the same 12 million token job, but assume an 80% cache hit rate, which is ordinary for a coding agent replaying its context. V4-Flash went from $0.8624 to $1.8160 off-peak, a 2.11x rise against the 1.80x on the uncached version of the same job. V4-Pro went from $2.6390 to $5.4560, a 2.07x rise. Cache your prompts well and you get a bigger increase, not a smaller one. That is the reverse of how nearly every other price rise on this blog has behaved.

One detail cuts the other way and is worth knowing. Because off-peak is defined as exactly half of peak on every line, the ratio between a cache hit and a cache miss is identical in both windows: 3.182% on Flash and 3.333% on Pro. Your cache discount does not get worse during peak hours. Only everything else does.

Both cards were quietly rebuilt on a 3:1 ratio

Underneath the multipliers there is a tidier structure than the six unmatched figures suggest. DeepSeek has been selling output at exactly twice input since V4 launched in April: $0.14 and $0.28, $0.435 and $0.87. From today both models sell output at exactly three times input: $0.22 and $0.66, $0.66 and $1.98. The Pro premium over Flash was 3.107x on both input and output and is now exactly 3.00x on both.

A 2:1 output ratio was always the outlier. OpenAI runs 6:1 on gpt-5.6-luna, Anthropic runs 5:1 across the Claude line, Google runs 5:1 on Gemini 3.7 Flash. DeepSeek moving to 3:1 is a step toward the industry shape rather than away from it, and it tells you which workloads got repriced: anything output-heavy. If your job is summarisation, extraction or classification, where output is a small fraction of input, you are near the 1.57x end of this. If it is code generation or long-form reasoning, you are near the 2.36x end.

Scoring the estimate we published eight days ago

On August 8, with nothing to go on but the word significant, we measured the room DeepSeek had rather than guessing the number. The claim was that it could raise its entire card 2.14x before anything comparable undercut it, with gpt-5.6-luna sitting just above that at 2.24x on the same job. That is now checkable, so here is the scorecard rather than a victory lap.

Off-peak came in at 1.80x, comfortably inside the room. Peak came in at 3.59x, well outside it. A workload spread evenly across the clock pays a blended 2.32x, which lands just past the line. So the honest reading is that the estimate was the right shape and the wrong single number, because the premise it rested on was that DeepSeek would publish one figure. It published two, and put the cheap one where most of the world's working hours are.

The same post also offered a precedent: when DeepSeek ended its V3 promotion in February 2025 it raised output 4x, input on a cache miss 2x and cache hits 5x. Measured against today's peak card, output moved 4.71x, input 3.14x, and cache hits on V4-Flash moved 5.000x. The cache multiplier repeated to three decimal places. That is either a house rule about what a cache read is worth relative to a miss, or a coincidence with a very narrow escape route.

Whether you ever pay peak is a question about your timezone

The windows are 01:00 to 04:00 and 06:00 to 10:00 UTC. DeepSeek states them in UTC now, which reads like a neutral international choice until you convert them back: 01:00 to 04:00 UTC is 09:00 to 12:00 in Beijing and 06:00 to 10:00 UTC is 14:00 to 18:00. These are the same two windows DeepSeek trailed at the end of June, put on the pricing page only in the first days of August, and pulled again by August 9 without ever charging them. The windows survived a round trip. Only the rates attached to them changed.

Local timeFirst windowSecond windowHits a 9 to 6 workday?
US Pacific, UTC-718:00 to 21:0023:00 to 03:00No
US Eastern, UTC-421:00 to 00:0002:00 to 06:00No
London, UTC+102:00 to 05:0007:00 to 11:00Yes, two hours of the morning
Berlin, UTC+203:00 to 06:0008:00 to 12:00Yes, three hours of the morning
India, UTC+5:3006:30 to 09:3011:30 to 15:30Yes, four and a half hours across midday
Beijing, UTC+809:00 to 12:0014:00 to 18:00Yes, both windows in full

An American team calling DeepSeek only while its own engineers are awake will never once touch the peak rate, and for them today's change is a flat 1.80x on the uncached job. A Berlin team loses its whole morning to it. A Bangalore team loses the middle of its day. The gap between the best and worst case is a factor of two on an identical workload, decided entirely by geography, and there is no setting or tier that adjusts it.

For anything running unattended the geography stops helping. Seven of twenty-four hours is 29.2% of the clock, so a cron job, a queue worker or an always-on agent pays the blended figure whether its owners are asleep or not. On the reference job that is $4.55 against $1.96 yesterday. Scale it to a month of 400M input and 80M output tokens and the bill goes from $78.40 to $181.87, or to $140.80 if you can keep every call out of those seven hours.

gpt-5.6-luna is now cheaper for anyone whose agents do not sleep

OpenAI cut Luna 80% on July 30 to $0.20 and $1.20. DeepSeek raises Flash today. Seventeen days apart, with nothing on the record linking them, and between them they have inverted a comparison that has held all year.

Route for 10M in, 2M outBillAgainst DeepSeek off-peak
V4-Flash weights, cheapest third-party host$0.940.27x
DeepSeek V4-Flash, off-peak$3.521.00x
gpt-5.6-luna$4.401.25x
DeepSeek V4-Flash, blended across 24h$4.551.29x
DeepSeek V4-Flash, peak$7.042.00x
DeepSeek V4-Pro, off-peak$10.563.00x
DeepSeek V4-Pro, peak$21.126.00x

Read the middle three rows together, because they are the decision. A team that controls when its calls happen keeps a 20% advantage over Luna. A team that does not gives it up and pays 3.3% more. The rate card no longer answers the question on its own, which is a new thing to have to say about DeepSeek, and it is the actual product of this change rather than any single multiplier.

The comparison narrows further once caching enters. On the 80% cached version of the job DeepSeek V4-Flash billed $0.8624 against Luna's $2.96, a 3.43x advantage. Off-peak that becomes $1.8160 against $2.96, or 1.63x. Blended it is $2.3457, or 1.26x. At peak DeepSeek bills $3.6320 and loses outright. Caching used to be the thing that made DeepSeek untouchable on price; it is now the thing that keeps it merely ahead. You can put your own token mix through the cost calculator rather than take a 5:1 reference job as representative, because the ratio moves this table around more than the rates do.

The strangest line in that table is the first one

DeepSeek-V4-Flash-0731 is MIT licensed and ungated on Hugging Face, so anyone can serve it. On August 16 the cheapest route on OpenRouter listed the same 0731 build at $0.0671 input and $0.1341 output, which prices the reference job at $0.94. DeepSeek's own off-peak card prices it at $3.52 and its peak card at $7.04. After today, the most expensive place to buy DeepSeek V4-Flash is DeepSeek, by 3.7x off-peak and 7.5x at peak against the cheapest route. Against the dearest third-party route the margin is nearer 9%, so the spread across hosts is wide and the cheapest number is the one doing the work in that comparison.

That gap existed before today at 2.09x, so this is a widening rather than a reversal, and the usual caveats apply: third-party routes vary in throughput, context handling and uptime, prompt caching is not implemented identically across hosts, and a listed price on an aggregator is not a contract. But it does mean this particular price rise is one of the few you can decline without changing model, rewriting prompts or re-running evals. Most price rises make you choose between your bill and your output quality. This one mostly asks who you want to pay.

On the Pro side there is a second event three days earlier that most coverage has treated separately. DeepSeek's changelog dates the general availability release of V4-Pro to August 13, and the pricing page now names the build DeepSeek-V4-Pro-0813. The model ID did not change, so anyone calling deepseek-v4-pro was moved onto it without doing anything. DeepSeek published self-reported numbers with it that are not modest: Terminal Bench 2.1 at 87.9, HLE at 42.7 without tools and 60.0 with them, DeepSWE 62.7, Cybergym 83.3, NL2Repo 61.5. A better model arriving three days before a dearer card is a defensible sequence and probably the intended one. It is worth noting anyway that the $0.435 being replaced was not a promotional rate DeepSeek was always going to withdraw. DeepSeek shipped Pro at 75% off in April, then in May cancelled the step-up that had been scheduled for May 31 and dropped the promotional label, describing the lower price as permanent. Permanent lasted twelve weeks. Current rates for every model named here sit on the pricing page.

Where every number here comes from

  • DeepSeek: Models and Pricing - Every DeepSeek rate in this post. The outgoing card at $0.14, $0.0028 and $0.28 for V4-Flash and $0.435, $0.003625 and $0.87 for V4-Pro, and footnote one carrying the replacement: off-peak at half of peak, peak defined as 01:00 to 04:00 and 06:00 to 10:00 UTC, effective 16:00 UTC on August 16, 2026, with the six new rates in a table. Also the model version strings DeepSeek-V4-Flash-0731 and DeepSeek-V4-Pro-0813, the 1M context window, the 384K maximum output and the concurrency limits of 2,500 and 500
  • DeepSeek: Change Log - The August 13, 2026 entry recording the GA release of DeepSeek-V4-Pro across app, web and API on the unchanged deepseek-v4-pro model name, along with the self-reported benchmark set quoted above: HLE 42.7 without tools and 60.0 with, Terminal Bench 2.1 87.9, NL2Repo 61.5, Cybergym 83.3, DeepSWE 62.7, Toolathlon-Verified 74.1, Agents' Last Exam 25.7, AutomationBench 31.8, DSBench-FullStack 71.1 and DSBench-Hard 67.2. None of these are independently verified and DeepSeek names no comparison baseline for them
  • TokenCost: DeepSeek makes the V4-Pro price cut permanent - Our May coverage of the 75% cut, the cancellation of the May 31 step-up to $1.74 and $3.48, and DeepSeek's own description of $0.435 and $0.87 as permanent rather than promotional. That is the claim today's card retires after twelve weeks
  • OpenAI: GPT-5.6 Luna - $0.20 input, $0.02 cached input and $1.20 output per million tokens, a 1,050,000 token context window and a February 16, 2026 knowledge cutoff. These are the figures behind the $4.40 and $2.96 comparison rows
  • OpenRouter: models endpoint - Read on August 16, 2026. The cheapest third-party route for the 0731 build at $0.0671 and $0.1341, which is the $0.94 row. Mind the slug: OpenRouter's bare deepseek/deepseek-v4-flash is the older 0423 build and cheaper again at $0.0615 and $0.1229, which is not the model DeepSeek is repricing. Note also that routed prices had not caught up to the change at the time of reading, with deepseek/deepseek-v4-pro-0813 still at the outgoing $0.435 and $0.87 and the undated deepseek/deepseek-v4-pro at $1.168 and $2.336, matching neither card. Re-check rather than assume
  • TokenCost: DeepSeek says a significant price increase is coming - Our August 8 post, the source of the 2.14x headroom estimate and the 2.24x gpt-5.6-luna line being scored above, and of the February 2025 V3 precedent of 4x output, 2x input and 5x cache hits
  • TokenCost: DeepSeek V4 will charge double during its busiest hours - The June 30 post covering the original announcement of these same windows as 09:00 to 12:00 and 14:00 to 18:00 Beijing time. Those rows were removed from both the English and Chinese pricing pages in early August without ever being charged, which is what makes today the first time DeepSeek has actually billed by the clock
  • Two things this post assumes rather than knows, stated plainly. The blended 2.32x figure and the $4.55 and $181.87 bills built on it assume traffic spread evenly across all 24 hours, which is true of a cron job and false of almost everything else; your own blend depends on your schedule and is the one number here you should compute rather than borrow. And DeepSeek has published no rationale for the split beyond allocating resources more reasonably, no commitment to hold these rates, and no model card or comparison baseline for the 0813 build behind its ten self-reported scores, so nothing above should be read as a forecast of where this card sits next month