Skip to main content
TokenCost logoTokenCost
GuideJune 30, 2026·7 min read

DeepSeek V4 will charge double during its busiest hours. Whether you ever pay the surcharge comes down to your timezone.

On June 30 DeepSeek told users that when V4 launches in mid-July, two daily windows will bill at twice the standard rate. No other major LLM API charges by the clock, so this is new ground. The headline writes itself as "DeepSeek doubles its prices," and that is the part worth slowing down on, because the seven peak hours sit in Beijing time. Map them onto where you actually work and the surcharge either swallows your mornings or never touches you at all.

A lit clock tower glowing against a dark winter night, evoking time-of-day API pricing

Photo by Juan Encalada on Unsplash

Standard versus peak, per million tokens

TierInput (yuan)Output (yuan)Input (~USD)Output (~USD)
V4 Pro, standard¥3¥6$0.44$0.88
V4 Pro, peak¥6¥12$0.88$1.76
V4 Flash, standard¥1¥2$0.15$0.29
V4 Flash, peak¥2¥4$0.29$0.59

Cache-miss input shown. Yuan figures are the official rates reported at the June 30 announcement; cache-hit input is far cheaper (¥0.025 standard, ¥0.05 peak on Pro) and also doubles. USD columns are our conversion near 6.8 yuan per dollar and are not a published rate card. Off-peak prices are unchanged from current V4 pricing.

What is actually changing

DeepSeek calls it peak-valley pricing, and the structure is simple. Two windows each day are designated peak: 09:00 to 12:00 and 14:00 to 18:00 Beijing time. Inside them, every token costs twice what it costs outside them. The other 17 hours, the valley, keep the standard rate you would pay today. The doubling is uniform across cache-hit input, cache-miss input, and output, so there is no line item that escapes it. DeepSeek frames the goal as a way to "allocate resources more rationally and improve service stability," which is the polite phrasing of a GPU shortage during business hours in China.

This is not the company's first experiment with the clock. Through 2025 DeepSeek ran the opposite scheme, a nighttime discount that knocked 50 to 75 percent off its older models between 16:30 and 00:30 UTC. That was a carrot to pull traffic into quiet hours. The V4 plan is the same idea wearing a stick: instead of rewarding the night, it taxes the day. The schedule arrives with the official V4 release in mid-July, and DeepSeek says it will send email 24 hours ahead of any change so nobody wakes up to a doubled invoice.

One caveat on the numbers. DeepSeek's public dollar pricing page already lists V4 Pro and V4 Flash at their standard rates, but its Chinese-language page still shows the old deepseek-chat and deepseek-reasoner models and has not yet posted the peak schedule. The peak figures here come from the June 30 announcement as carried by Odaily and others, not from a live rate card, so treat them as firm but freshly reported.

Where the peak hours land in your day

Here is the detail that decides whether any of this matters to you. The windows are fixed to Beijing time, not yours. Convert them and the surcharge falls in wildly different places depending on where your servers and your engineers sit. The two Beijing windows are 01:00 to 04:00 and 06:00 to 10:00 UTC. Translated into local working hours, late June:

RegionPeak windows, localHits a 9-to-5?
US Pacific (PDT)6pm-9pm, 11pm-3amNo
US Eastern (EDT)9pm-12am, 2am-6amNo
UK (BST)2am-5am, 7am-11amPartly
Central Europe (CEST)3am-6am, 8am-12pmPartly
India (IST)6:30am-9:30am, 11:30am-3:30pmYes
Beijing / Singapore9am-12pm, 2pm-6pmYes
Tokyo (JST)10am-1pm, 3pm-7pmYes

Read down the last column and the pattern is hard to miss. For anyone on US time the entire surcharge falls in the evening and overnight. A San Francisco or New York team that calls DeepSeek between breakfast and dinner is in the valley the whole time and pays the standard rate, full stop. The people who actually eat the doubling are in Asia, where the windows were drawn to begin with, and in Europe, where the second window clips the morning. A team in Frankfurt or London doing real work at 9am is paying peak rates for a chunk of every morning.

There is a quiet irony here for US shops that picked DeepSeek partly to dodge American frontier prices. The geography that makes the model feel like a foreign dependency is the same geography that hands you its cheapest hours. Beijing's busy afternoon is your quiet night.

"Double" is the worst case, not the average

The word doubling does a lot of scary work in the headlines, and it is only true for tokens spent entirely inside the peak windows. Seven of the day's 24 hours carry the surcharge. A workload smeared evenly across the clock, the shape of a 24/7 agent or a global product with users in every timezone, spends 17 hours at 1x and seven at 2x. That averages to 31/24, or about 29 percent over the standard bill. Real, but a long way from double. Take a V4 Pro agent burning 50M input and 10M output tokens a month:

How the work is timedMonthly costvs standard
All off-peak (US workday or scheduled)$30.45baseline
Spread evenly across 24 hours$39.33+29%
All inside peak windows$60.90+100%

The gap between the top and bottom rows is the whole game. Same model, same token count, and a 2x spread in the bill decided entirely by when the requests go out. For interactive traffic you cannot move, your timezone sets where in that range you land. For batch work you can move, the schedule is a dial you control, and most teams have never had a reason to touch it before now.

What to do about it

If you are on US time and your traffic follows your working day, do nothing. You are already in the valley and the surcharge will not find you. The one thing worth checking is your cron jobs: nightly batch runs, evening backfills, and scheduled evaluations on US time can drift straight into the Beijing afternoon, so move anything heavy and non-urgent to your own daytime, which is DeepSeek's night.

For European and Asian teams the lever is the morning. The second peak window is the expensive one because it overlaps the start of the workday. Queue document processing, embeddings refreshes, and any bulk generation for the early afternoon valley instead of running it the moment people log on. If your jobs are latency-tolerant, a scheduler that simply avoids 06:00 to 10:00 UTC erases most of the surcharge without touching a single interactive request.

And do not forget the cache. Peak doubling applies to cache-hit input too, but the absolute number is so small, half a cent per million on Pro even at peak, that aggressive context caching still flattens your bill far more than chasing the clock ever will. The surge is a timing problem worth a scheduler rule; it is not a reason to abandon a model that, off-peak, still undercuts almost everything in its class. We keep the standard rates current on the pricing page and you can model your own mix in the calculator.

The part that outlasts this one model

Strip away the V4 specifics and what DeepSeek just did is import a pricing idea from electricity and airlines into the LLM market, apparently the first major provider to try it. Inference is a capacity business with a brutal daily demand curve, so time-of-day pricing was always going to show up eventually. If it works for DeepSeek, expect the cheaper Chinese labs that compete on price to copy it within a quarter, and watch whether the Western providers, who have leaned on flat rates and opaque priority tiers instead, feel any pressure to follow.

For now the practical takeaway is small and concrete. Your DeepSeek bill is about to depend on a clock set in another country. Find out where those seven hours fall in your day, move what you can out of them, and stop reading the word "double" as if it applies to you. For most teams reading this in dollars, it does not.

Sources