Skip to main content
TokenCost logoTokenCost

DeepSeek V4.1 Flash

Last verified September 10, 2026 · DeepSeek pricing

$0.150/1M input · $0.600/1M output · 1.0M context · DeepSeek

Count Tokens

Tokens
Input Cost
Output Cost

Estimate Monthly Cost

Monthly Cost Estimator

Quick:
<$0.0001/mo

Pick a preset above or enter custom usage

Alternatives to DeepSeek V4.1 Flash

Pricing Details

DeepSeekModel name is deepseek-flash (no version in the id), version DeepSeek-V4.1-Flash, released 2026-09-10 with prices effective 04:00 UTC that day. Card: off-peak $0.15 input (cache miss) / $0.003 cache hit / $0.60 output; peak exactly double at $0.30 / $0.006 / $1.20. Peak is 01:00-04:00 and 06:00-10:00 UTC, MONDAY THROUGH FRIDAY (35 of 168 hours a week, 20.8%); the stored figure is the OFF-PEAK rate. A workload spread evenly across the week pays 1.2083x off-peak: blended $0.18125 input / $0.725 output / $0.003625 cache hit. Against the V4 Flash (0731) card it retires ($0.22 / $0.007 / $0.66) the cuts are 31.8% input, 57.1% cache hit, 9.1% output; the 60%/33%/11% figures in some coverage are the yuan card (¥0.05/¥1.5/¥4.5 to ¥0.02/¥1/¥4), which was set at 6.67 yuan per dollar against the 6.82 every other line on the page still implies. Cache hit is 2% of input, restoring the 98% discount V4 Flash launched with in April. Native vision (JPEG/PNG/GIF/WebP), thinking on by default at effort high, effort low/high/max, Responses API and Anthropic-format endpoint supported, concurrency 2,500 per user_id. The legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp still resolve but are 'temporarily routed' to this model and billed at this card, with no stated end to the routing and no snapshot of the old model. FROM 2026-09-14 12:00 BEIJING (04:00 UTC) deepseek-v4-pro IS ALSO ROUTED HERE and billed at this card 'until V4.1 Pro is released in the future'; no V4.1 Pro date exists. Architecture per the V4.1 technical report: 552B backbone parameters, causal encoder-decoder, 16B active per token during decode and 8B during prefill, 384 routed experts with 6 active, 40 layers, FP8 with FP4 experts, KV cache 890 bytes per token (about a quarter of V4 Flash), pretrained on 45T multimodal tokens. Weights on Hugging Face under MIT; the README was titled DeepSeek-V4.1-Exp for its first two hours, renamed 06:25 UTC 2026-09-10. DeepSeek-reported scores from the 2026-09-10 changelog: GPQA Diamond 90.9, HLE 36.8 (63.9 with tools), Codeforces 3471, MathArena Apex 65.6, Terminal-Bench 2.1 90.6, Terminal-Bench 3.0 30.0, Terminal-Bench 4.0 31.2, DeepSWE v1.1 74.2, NL2Repo 65.4, CyberGym 88.1, Automation-Bench 54.8 (V4 Pro 43.2 on the same Pass@1 row), Agents' Last Exam 31.8. Of the sixteen with a V4 Pro (0813) counterpart in Table 3 of the V4.1 tech report, Pro is still higher on two: GPQA Diamond (92.4) and HLE without tools (42.7 vs 39.1 on the text-only subset). NO ARTIFICIAL ANALYSIS ENTRY as of 2026-09-10; AA's v4.3 index has V4 Flash 0731 at 35 and V4 Pro 0813 at 36. Throughput figures circulating (355-427 tok/s, 178 ms TTFT) are individual users' tests and appear in no DeepSeek document. Shares the $0.15 input price with GLM-5.3-Flash and Qwen3.8-Flash; dearest of the three on output ($0.60 vs $0.50 and $0.47) and cheapest by far on cache hits ($0.003 vs $0.03 and $0.015).
Input / 1M tokens
$0.15
Output / 1M tokens
$0.6
Context Window
1.0M
Max Output
384K

Price History

Launched at current price on 2026-09-10. No price changes recorded since.

Frequently Asked Questions

Common questions about DeepSeek V4.1 Flash pricing and usage

Read More About DeepSeek V4.1 Flash