Skip to main content
TokenCost logoTokenCost

Qwen3.8 Flash

Last verified September 10, 2026 · Alibaba pricing ↗

$0.150/1M input · $0.470/1M output · 1.0M context · Alibaba

Count Tokens

—
Tokens
—
Input Cost
—
Output Cost

Estimate Monthly Cost

Monthly Cost Estimator

Quick:
<$0.0001/mo

Pick a preset above or enter custom usage

Alternatives to Qwen3.8 Flash

Pricing Details

AlibabaCORRECTED 2026-09-10: Alibaba Cloud Model Studio's own pricing page lists qwen3.8-flash (International) at $0.15 input / $0.47 output per 1M, a single tier from 0 to 1M tokens, with cache hits at 10% of input ($0.015) and explicit cache creation at 125%. The $0.16 previously stored came from Alibaba's launch blog's rounding of the yuan card and from OpenRouter; the page a card is charged against says $0.15, which puts this model on the same input price as GLM-5.3-Flash and DeepSeek V4.1 Flash. The rest of this note predates the correction and still says $0.16 where it discusses the dollar card. TWO SEPARATE RATE CARDS, not one converted: mainland China meters this model at 1 yuan per million input and 3 yuan per million output and is reported 60-70% cheaper, while $0.16/$0.47 is the international (Singapore) tier's own USD card. Do NOT describe the dollar figure as an FX conversion of the yuan one - the arithmetic refuses it. 1 yuan is about $0.14 at recent rates, not $0.16, and the yuan card's output:input ratio is exactly 3.0 against the dollar card's 2.94. So the USD rate does not drift with FX; which card you pay is a function of which endpoint you call. We could not reach a first-party Alibaba Model Studio page carrying the USD rate, hence priceUnverified; OpenRouter's Alibaba endpoint is the source. Cached input $0.016, a 10x discount off input, against Z.ai's 5x on GLM-5.3-Flash. CACHE WRITE IS $0.20, i.e. 1.25x its own fresh input rate. The write REPLACES the fresh input charge rather than adding to it, so a prefix written and never reused is worse than no cache by exactly $0.04/Mtok, and break-even is 0.28 of a read: writing then reading once costs $0.216 against $0.32 for sending the same tokens twice uncached. Any prefix used even once is ahead. Do NOT compute this as 0.20/(0.16-0.016)=1.389, which wrongly treats the write as additive. supports_implicit_caching TRUE, so prefixes match automatically - a team that writes no caching code gets Alibaba's rate and not Z.ai's. Listed on OpenRouter 2026-08-26 at 19:37 UTC, 5h38m after GLM-5.3-Flash, which is its closest comparable on price, modality, context and target workload. 125B total parameters plus a 51B N-gram embedding layer, 6B active per token, multimodal MoE, text+image+video in. Native context 262,144 extended to 1M via YaRN. REASONING IS OPTIONAL: the listing reports mandatory false with reasoning-budget support, so output volume (and therefore the larger half of the bill) is cappable, which GLM-5.3-Flash's mandatory max-effort configuration does not allow. NO ARTIFICIAL ANALYSIS ENTRY as of 2026-08-27, so there is no independent index score, token count or cost-per-task; every cross-model cost figure we publish holds GLM-5.3-Flash's measured tokens constant across both, which prices the rate cards and not the models. OpenRouter P50 telemetry: 52 output tok/s, 4.68s latency, 100% uptime, one endpoint (Alibaba Cloud International). Do not put that latency in a column beside AA's TTFT figures; different measurement. Vendor-reported benchmarks, unchecked: DeepSWE 1.1 58.7, SWE-bench Pro 62.5, CoWorkBench 73.9, GPQA Diamond 91.7, LiveCodeBench v6 91.9, AndroidWorld 84.5, LVBench 76.6. Separately, Alibaba open-weighted Qwen3.8-Flash-Next on 2026-08-24, whose native context is 262K rather than the 1M sold here; it is a sibling, not this model, and self-host-vs-API math is not like-for-like. Against Qwen3.7 Flash the upgrade is not uniformly a rise: 3.7's tiering charges $0.03/$0.13 under 32K but $0.20/$0.80 above 256K, so this flat card is 5.33x dearer on short prompts and 1.25x/1.70x CHEAPER on long ones. Estimated tokens (different tokenizer).
Input / 1M tokens
$0.15
Output / 1M tokens
$0.47
Context Window
1.0M
Max Output
131K

Price History

No price change recorded since 2026-08-27, when this site started tracking this model; it launched on 2026-08-26.

Frequently Asked Questions

Common questions about Qwen3.8 Flash pricing and usage