Qwen3.8 Max
Last verified August 4, 2026 · Alibaba pricing ↗$2.00/1M input · $6.00/1M output · 1.0M context · Alibaba
Count Tokens
—
Tokens
—
Input Cost
—
Output Cost
Estimate Monthly Cost
Monthly Cost Estimator
Quick:
<$0.0001/mo
Pick a preset above or enter custom usage
Alternatives to Qwen3.8 Max
Pricing Details
AlibabaClosed-weights multimodal flagship (text/image/video input, text output), 2.4T total / 95B active MoE. GA 2026-08-03 after a 2026-07-19 preview at WAIC. Price stored is the gateway consensus: OpenRouter, Vercel AI Gateway, NanoGPT, Kilo and Merge all quote $2.00/$6.00 with cache read $0.25 and cache write $2.50. Alibaba Model Studio publishes NO per-token rate for this model as of 2026-08-04 (its table still tops out at Qwen3.7 Max), so there is no first-party card to check against. Resellers apply flat margins to the same baseline: AIHubMix 0.845x on both legs, Deep Infra 0.825x but capped to 256K context, CrossModel 0.94x; Qubrid is the only independent card at $2.87/$7.12 list, 20% off at launch. Flat pricing across the full 1M window, no context tiering. Cache read of $0.25 is 12.5% of input and matches neither Alibaba's documented 10% explicit nor 20% implicit rate; it is exactly 10% of $2.50, which may mean $2 is itself a promo. Reasoning budget 262K, reasoning_effort low/medium/xhigh with xhigh default, thinking tokens billed as output. 991K usable input, ~983K with thinking on. AA Intelligence Index 53; generated 150M tokens running the index vs a 63M median (2.38x) at a cost of $2,159.51, output 46.5 tok/s, TTFT 2.48s. Against Kimi K3 ($2,437.41, 130M tokens, ~36 tok/s, index 57) the 60% output-price advantage yields only an 11.4% smaller bill. Beware secondary sources quoting $2,690.80 and 62 tok/s for Kimi K3; AA's own page says otherwise. Vendor-run benchmarks: Terminal-Bench 2.1 86.6, SWE-Bench Pro 67.7, PaperBench 93.0, IFBench 82.8, GPQA Diamond 92.6, HLE 43.6, FrontierSWE 73.5, OSWorld-Verified 86.1. Alibaba leads only two of six rows on its own table. Rate limits 2M TPM / 15K RPM. Not on Bedrock or Vertex. No published knowledge cutoff or model card; weights promised on HF/ModelScope but not shipped as of 2026-08-04. Two different overnight discounts get conflated: Alibaba documents a real per-token 80% night discount (22:00-08:00 UTC+8) for Qwen3.7 Max on HK/Frankfurt/US-Virginia, unconfirmed for 3.8; separately Qoder's GA credits promo is 50% off (0.5x to 0.25x multiplier), credits only, and its expired preview promo was ~90% all-day stacking to ~98% off-peak.
Input / 1M tokens
$2
Output / 1M tokens
$6
Context Window
1.0M
Max Output
131K
Price History
No price change recorded since 2026-08-04, when this site started tracking this model; it launched on 2026-08-03.
Frequently Asked Questions
Common questions about Qwen3.8 Max pricing and usage
Read More About Qwen3.8 Max
Sakana priced Fugu Max at $2 and $6, which is Qwen3.8 Max's card to the cent, and says the output line is 40% to 60% under Sonnet 5, Terra and Kimi K3. The tokens billed at that rate include every internal call the orchestrator makes, and Sakana publishes neither the count nor the models it calls.
September 12, 2026 · 14 min read
Alibaba says Qwen3.8 Max is the world's second-best model. It won't publish a benchmark, an active-parameter count, or a per-token price.
July 22, 2026 · 7 min read