Skip to main content
TokenCost logoTokenCost

DeepSeek V4-Flash

Last verified August 21, 2026 · DeepSeek pricing

$0.220/1M input · $0.660/1M output · 1.0M context · DeepSeek

Count Tokens

Tokens
Input Cost
Output Cost

Estimate Monthly Cost

Monthly Cost Estimator

Quick:
<$0.0001/mo

Pick a preset above or enter custom usage

Alternatives to DeepSeek V4-Flash

Pricing Details

DeepSeekPRICE CHANGE 2026-08-16: DeepSeek's warned-of "significant increase" landed as a peak/off-peak split effective 16:00 UTC on 2026-08-16. Peak is 01:00-04:00 and 06:00-10:00 UTC (7 of 24 hours); off-peak is exactly half of peak on every line. New card: off-peak $0.22 input / $0.66 output / $0.007 cache hit, peak $0.44 / $1.32 / $0.014. The stored figure is the OFF-PEAK rate, which applies 17 of 24 hours; an evenly spread workload pays a blended $0.2842 input and $0.8525 output. Multipliers off the old $0.14/$0.28/$0.0028 card are not uniform: input 1.57x off-peak, output 2.36x, cache hits 2.50x, and 3.14x/4.71x/5.00x at peak. Cache discount narrows from 98% to 96.8% of input. Output:input ratio moved from 2:1 to exactly 3:1. Weights are MIT and ungated, and third-party hosts undercut DeepSeek's own card: OpenRouter's cheapest deepseek-v4-flash route read $0.0615/$0.1229 on 2026-08-16, about a quarter of the new off-peak rate. 304.18B total params per the Hugging Face safetensors index (Artificial Analysis and most coverage say 284B; 13B active is widely repeated but DeepSeek states neither figure itself). Cache hit: $0.0028/1M, a 98% discount against the 90% most of the industry offers. Three reasoning modes. Replaces deepseek-chat and deepseek-reasoner aliases (deprecated 2026-07-24). Silently upgraded to the 0731 build on 2026-07-31 under the same model ID and the same rate card: architecture and size unchanged, "only re-post-trained" per DeepSeek's changelog, plus native Responses API support and Codex adaptation. Weights are MIT and ungated. Artificial Analysis Intelligence Index 50 (49.93), up 10 points from the April build, ranked #3 among 101 open-weights models (closed models are ranked against a larger field); $72.02 and ~206M output tokens to run the index, GPQA Diamond 91%, Terminal-Bench 2.1 79% (DeepSeek self-reports 82.7), AA-LCR 66%, GDPval-AA v2 1559. AA-Omniscience -16 with an 84% hallucination rate, which AA defines as incorrect answers as a share of responses that were not correct, so it does NOT mean 84% of all answers are wrong. Very verbose: 36th of 101 open-weights models on output volume against a 100M-token median for that cohort. Concurrency limit 2,500 against V4-Pro's 500. Peak-hour pricing was first announced 2026-06-30 and published on the pricing page through early August (2x during 09:00-12:00 and 14:00-18:00 Beijing time), then removed from both language versions around 2026-08-08 without ever being charged, and replaced by a notice of a coming "significant" increase with no number or date. It returned on 2026-08-16 with the same two windows restated in UTC and new rates attached, which is the change recorded at the top of this note. The old 16:30-00:30 UTC off-peak discount belonged to V3/R1 and no longer exists.
Input / 1M tokens
$0.22
Output / 1M tokens
$0.66
Context Window
1.0M
Max Output
384K

Price History

DateInput /1MOutput /1MChange
2026-08-21$0.140 $0.220$0.280 $0.660 +57%

Frequently Asked Questions

Common questions about DeepSeek V4-Flash pricing and usage

Read More About DeepSeek V4-Flash