Skip to main content
TokenCost logoTokenCost

Qwen3.8 Flash

Last verified August 27, 2026 · Alibaba pricing

$0.160/1M input · $0.470/1M output · 1.0M context · Alibaba

Count Tokens

Tokens
Input Cost
Output Cost

Estimate Monthly Cost

Monthly Cost Estimator

Quick:
<$0.0001/mo

Pick a preset above or enter custom usage

Alternatives to Qwen3.8 Flash

Pricing Details

AlibabaTWO SEPARATE RATE CARDS, not one converted: mainland China meters this model at 1 yuan per million input and 3 yuan per million output and is reported 60-70% cheaper, while $0.16/$0.47 is the international (Singapore) tier's own USD card. Do NOT describe the dollar figure as an FX conversion of the yuan one - the arithmetic refuses it. 1 yuan is about $0.14 at recent rates, not $0.16, and the yuan card's output:input ratio is exactly 3.0 against the dollar card's 2.94. So the USD rate does not drift with FX; which card you pay is a function of which endpoint you call. We could not reach a first-party Alibaba Model Studio page carrying the USD rate, hence priceUnverified; OpenRouter's Alibaba endpoint is the source. Cached input $0.016, a 10x discount off input, against Z.ai's 5x on GLM-5.3-Flash. CACHE WRITE IS $0.20, i.e. 1.25x its own fresh input rate. The write REPLACES the fresh input charge rather than adding to it, so a prefix written and never reused is worse than no cache by exactly $0.04/Mtok, and break-even is 0.28 of a read: writing then reading once costs $0.216 against $0.32 for sending the same tokens twice uncached. Any prefix used even once is ahead. Do NOT compute this as 0.20/(0.16-0.016)=1.389, which wrongly treats the write as additive. supports_implicit_caching TRUE, so prefixes match automatically - a team that writes no caching code gets Alibaba's rate and not Z.ai's. Listed on OpenRouter 2026-08-26 at 19:37 UTC, 5h38m after GLM-5.3-Flash, which is its closest comparable on price, modality, context and target workload. 125B total parameters plus a 51B N-gram embedding layer, 6B active per token, multimodal MoE, text+image+video in. Native context 262,144 extended to 1M via YaRN. REASONING IS OPTIONAL: the listing reports mandatory false with reasoning-budget support, so output volume (and therefore the larger half of the bill) is cappable, which GLM-5.3-Flash's mandatory max-effort configuration does not allow. NO ARTIFICIAL ANALYSIS ENTRY as of 2026-08-27, so there is no independent index score, token count or cost-per-task; every cross-model cost figure we publish holds GLM-5.3-Flash's measured tokens constant across both, which prices the rate cards and not the models. OpenRouter P50 telemetry: 52 output tok/s, 4.68s latency, 100% uptime, one endpoint (Alibaba Cloud International). Do not put that latency in a column beside AA's TTFT figures; different measurement. Vendor-reported benchmarks, unchecked: DeepSWE 1.1 58.7, SWE-bench Pro 62.5, CoWorkBench 73.9, GPQA Diamond 91.7, LiveCodeBench v6 91.9, AndroidWorld 84.5, LVBench 76.6. Separately, Alibaba open-weighted Qwen3.8-Flash-Next on 2026-08-24, whose native context is 262K rather than the 1M sold here; it is a sibling, not this model, and self-host-vs-API math is not like-for-like. Against Qwen3.7 Flash the upgrade is not uniformly a rise: 3.7's tiering charges $0.03/$0.13 under 32K but $0.20/$0.80 above 256K, so this flat card is 5.33x dearer on short prompts and 1.25x/1.70x CHEAPER on long ones. Estimated tokens (different tokenizer).
Input / 1M tokens
$0.16
Output / 1M tokens
$0.47
Context Window
1.0M
Max Output
131K

Unverified rate. Alibaba has not published a price for Qwen3.8 Flash. Treat these figures as indicative and confirm with Alibaba before relying on them.

Price History

Launched at current price on 2026-08-26. No price changes recorded since.

Frequently Asked Questions

Common questions about Qwen3.8 Flash pricing and usage