DeepSeek V4-Flash
Last verified August 21, 2026 · DeepSeek pricing ↗$0.220/1M input · $0.660/1M output · 1.0M context · DeepSeek
Count Tokens
—
Tokens
—
Input Cost
—
Output Cost
Estimate Monthly Cost
Monthly Cost Estimator
Quick:
<$0.0001/mo
Pick a preset above or enter custom usage
Alternatives to DeepSeek V4-Flash
Pricing Details
DeepSeekPRICE CHANGE 2026-08-16: DeepSeek's warned-of "significant increase" landed as a peak/off-peak split effective 16:00 UTC on 2026-08-16. Peak is 01:00-04:00 and 06:00-10:00 UTC (7 of 24 hours); off-peak is exactly half of peak on every line. New card: off-peak $0.22 input / $0.66 output / $0.007 cache hit, peak $0.44 / $1.32 / $0.014. The stored figure is the OFF-PEAK rate, which applies 17 of 24 hours; an evenly spread workload pays a blended $0.2842 input and $0.8525 output. Multipliers off the old $0.14/$0.28/$0.0028 card are not uniform: input 1.57x off-peak, output 2.36x, cache hits 2.50x, and 3.14x/4.71x/5.00x at peak. Cache discount narrows from 98% to 96.8% of input. Output:input ratio moved from 2:1 to exactly 3:1. Weights are MIT and ungated, and third-party hosts undercut DeepSeek's own card: OpenRouter's cheapest deepseek-v4-flash route read $0.0615/$0.1229 on 2026-08-16, about a quarter of the new off-peak rate. 304.18B total params per the Hugging Face safetensors index (Artificial Analysis and most coverage say 284B; 13B active is widely repeated but DeepSeek states neither figure itself). Cache hit: $0.0028/1M, a 98% discount against the 90% most of the industry offers. Three reasoning modes. Replaces deepseek-chat and deepseek-reasoner aliases (deprecated 2026-07-24). Silently upgraded to the 0731 build on 2026-07-31 under the same model ID and the same rate card: architecture and size unchanged, "only re-post-trained" per DeepSeek's changelog, plus native Responses API support and Codex adaptation. Weights are MIT and ungated. Artificial Analysis Intelligence Index 50 (49.93), up 10 points from the April build, ranked #3 among 101 open-weights models (closed models are ranked against a larger field); $72.02 and ~206M output tokens to run the index, GPQA Diamond 91%, Terminal-Bench 2.1 79% (DeepSeek self-reports 82.7), AA-LCR 66%, GDPval-AA v2 1559. AA-Omniscience -16 with an 84% hallucination rate, which AA defines as incorrect answers as a share of responses that were not correct, so it does NOT mean 84% of all answers are wrong. Very verbose: 36th of 101 open-weights models on output volume against a 100M-token median for that cohort. Concurrency limit 2,500 against V4-Pro's 500. Peak-hour pricing was first announced 2026-06-30 and published on the pricing page through early August (2x during 09:00-12:00 and 14:00-18:00 Beijing time), then removed from both language versions around 2026-08-08 without ever being charged, and replaced by a notice of a coming "significant" increase with no number or date. It returned on 2026-08-16 with the same two windows restated in UTC and new rates attached, which is the change recorded at the top of this note. The old 16:30-00:30 UTC off-peak discount belonged to V3/R1 and no longer exists.
Input / 1M tokens
$0.22
Output / 1M tokens
$0.66
Context Window
1.0M
Max Output
384K
Price History
| Date | Input /1M | Output /1M | Change |
|---|---|---|---|
| 2026-08-21 | $0.140 → $0.220 | $0.280 → $0.660 | ↑ +57% |
Frequently Asked Questions
Common questions about DeepSeek V4-Flash pricing and usage
Read More About DeepSeek V4-Flash
DeepSeek said the increase would be significant. It is 1.80x off-peak and 3.59x at peak, and only one of those two numbers fits inside the room we measured eight days ago.
August 16, 2026 · 10 min read
DeepSeek kept the price at $0.14 and added ten index points. Eleven models now tie on intelligence, and the bill for proving it runs from $72 to $1,061.
August 2, 2026 · 10 min read
DeepSeek switched off deepseek-chat and deepseek-reasoner yesterday. The rename is one line; the reasoning default it flips on is what moves your bill.
July 25, 2026 · 8 min read