Skip to main content
TokenCost logoTokenCost

Qwen3.8 Max

Last verified August 4, 2026 · Alibaba pricing

$2.00/1M input · $6.00/1M output · 1.0M context · Alibaba

Count Tokens

Tokens
Input Cost
Output Cost

Estimate Monthly Cost

Monthly Cost Estimator

Quick:
<$0.0001/mo

Pick a preset above or enter custom usage

Alternatives to Qwen3.8 Max

Pricing Details

AlibabaClosed-weights multimodal flagship (text/image/video input, text output), 2.4T total / 95B active MoE. GA 2026-08-03 after a 2026-07-19 preview at WAIC. Price stored is the gateway consensus: OpenRouter, Vercel AI Gateway, NanoGPT, Kilo and Merge all quote $2.00/$6.00 with cache read $0.25 and cache write $2.50. Alibaba Model Studio publishes NO per-token rate for this model as of 2026-08-04 (its table still tops out at Qwen3.7 Max), so there is no first-party card to check against. Resellers apply flat margins to the same baseline: AIHubMix 0.845x on both legs, Deep Infra 0.825x but capped to 256K context, CrossModel 0.94x; Qubrid is the only independent card at $2.87/$7.12 list, 20% off at launch. Flat pricing across the full 1M window, no context tiering. Cache read of $0.25 is 12.5% of input and matches neither Alibaba's documented 10% explicit nor 20% implicit rate; it is exactly 10% of $2.50, which may mean $2 is itself a promo. Reasoning budget 262K, reasoning_effort low/medium/xhigh with xhigh default, thinking tokens billed as output. 991K usable input, ~983K with thinking on. AA Intelligence Index 53; generated 150M tokens running the index vs a 63M median (2.38x) at a cost of $2,159.51, output 46.5 tok/s, TTFT 2.48s. Against Kimi K3 ($2,437.41, 130M tokens, ~36 tok/s, index 57) the 60% output-price advantage yields only an 11.4% smaller bill. Beware secondary sources quoting $2,690.80 and 62 tok/s for Kimi K3; AA's own page says otherwise. Vendor-run benchmarks: Terminal-Bench 2.1 86.6, SWE-Bench Pro 67.7, PaperBench 93.0, IFBench 82.8, GPQA Diamond 92.6, HLE 43.6, FrontierSWE 73.5, OSWorld-Verified 86.1. Alibaba leads only two of six rows on its own table. Rate limits 2M TPM / 15K RPM. Not on Bedrock or Vertex. No published knowledge cutoff or model card; weights promised on HF/ModelScope but not shipped as of 2026-08-04. Two different overnight discounts get conflated: Alibaba documents a real per-token 80% night discount (22:00-08:00 UTC+8) for Qwen3.7 Max on HK/Frankfurt/US-Virginia, unconfirmed for 3.8; separately Qoder's GA credits promo is 50% off (0.5x to 0.25x multiplier), credits only, and its expired preview promo was ~90% all-day stacking to ~98% off-peak.
Input / 1M tokens
$2
Output / 1M tokens
$6
Context Window
1.0M
Max Output
131K

Price History

Launched at current price on 2026-08-03. No price changes recorded.

Frequently Asked Questions

Common questions about Qwen3.8 Max pricing and usage

Read More About Qwen3.8 Max