Skip to main content
TokenCost logoTokenCost
Model ReleaseJuly 22, 2026·7 min read

Alibaba says Qwen3.8 Max is the world's second-best model. It won't publish a benchmark, an active-parameter count, or a per-token price.

Qwen3.8-Max-Preview showed up at WAIC in Shanghai around July 19, two days after Moonshot open-weighted Kimi K3. Alibaba wrapped it in a big number, 2.4 trillion parameters, and a bigger claim, second only to Claude Fable 5. What it did not ship is the part a buyer needs: no benchmark sheet, no active-parameter count, and no per-token rate. You cannot even buy tokens for it in the normal sense. The only meter is a credit subscription, and Alibaba never says what a credit is worth in tokens.

Aerial night view of the Shanghai skyline, host city of WAIC 2026 where Qwen3.8 Max was previewed

Photo by Zhou Xian on Unsplash

A flagship with a price tag you cannot read

Start with the one thing that matters for anyone budgeting a model: what does a million tokens cost. For Qwen3.8 Max the honest answer is that nobody outside Alibaba knows, because Alibaba has not published a per-token rate. Every other Max-tier flagship this year came with a rate card on day one. Fable 5 arrived at $10 and $50. GPT-5.6 Sol arrived at $5 and $30. Kimi K3 arrived at $3 and $15. Qwen3.8 Max arrived with a subscription page.

Access runs through Alibaba's Token Plan, a credit-based subscription, plus the Qoder and QoderWork coding products. During the preview, Alibaba says the meter runs at about 10% of standard pricing. That sounds like a discount, and it is, but a discount off an unpublished number is not a price. You are being asked to commit to a monthly plan for a model whose unit economics are not on the page.

This matters more than a missing spec line. A pricing tool exists to answer one question, is this model cheaper than that one for my workload, and Qwen3.8 Max is currently unanswerable. So the rest of this post does two things. It decodes the plan you can actually sign up for, and it anchors what a token is likely to cost when Alibaba finally prints a rate, using the model it replaces as the reference point.

The Token Plan, decoded

Here is what you can buy today. The Token Plan sells credits, not tokens, in three individual tiers. Each grants a bucket of credits over a rolling window and caps how many agents you can run at once. Dollar figures below are conversions from the yuan list, so they wobble a few dollars depending on the source and the day.

TierApprox / monthCredits / 7 daysConcurrent agents
Lite~$62,5001–2
Standard~$18–2010,0003–4
Pro~$68–7440,0006–8

Now the discount stack, which is where it gets genuinely interesting. Preview access already runs at roughly 10% of standard rates. On top of that, credit consumption carries a 90% daytime discount at launch, and individual users get a further 80% off at night, between 22:00 and 08:00 Beijing time. Stack the preview rate against the overnight cut and you land near 2% of the original price for off-peak work. That is not a typo. Alibaba is paying you, in effect, to move batch jobs to the small hours.

And here is the catch that undoes all of it as a comparison. The plans quote credits. The one number that would turn credits into a token cost, how many tokens a single credit buys on Qwen3.8 Max, appears on neither pricing page, as Digital Applied's teardown of the Token Plan also found. A Lite plan gives you 2,500 credits a week, but 2,500 credits could be a handful of long agent runs or a lot of short ones depending on a conversion Alibaba keeps to itself. Until that rate is published, any per-token figure you see for this model, including one you could back out of these plans, is a guess wearing a decimal point.

"Second only to Fable 5," with no scoresheet

The capability pitch has the same shape as the pricing: a bold headline, no supporting data. Alibaba's Qwen account positioned the model as comparable to leading frontier systems and second only to Claude Fable 5. Read that as marketing, because it is not a benchmark. There is no Artificial Analysis Intelligence Index entry, no LMArena placement, no task-level table, not even a comparison against Alibaba's own Qwen3.7 Max. For a 2.4 trillion parameter model claiming the number-two slot on Earth, the evidence on the table is a sentence.

Tier-plausible is the fair read, which is why the claim is not worth dismissing outright. The predecessor set a real bar. Qwen3.7 Max scored 56.6 on the AA Intelligence Index, the top Chinese model at the time, with GPQA-Diamond around 92 and SWE-Bench Verified in the low 80s. A genuine step up from that would put Qwen3.8 Max in the same conversation as Kimi K3, which sits at 57 on the same index. Whether it clears Fable 5's neighborhood is exactly the thing a benchmark would tell you, and exactly the thing Alibaba withheld.

The specs are half-drawn too. The 2.4T headline is a total parameter count for what is described as a sparse mixture-of-experts model, but the active-parameter count, the number that actually drives inference cost, is not disclosed. The 1M context window everyone is quoting is inherited from tooling and the 3.7 Max spec, not confirmed on a 3.8 Max sheet. Max output, knowledge cutoff, and the open-weight license are all blank. Multimodality is the one concrete addition: Alibaba says this is the first Qwen above a trillion parameters to take images, video, and documents alongside text.

What it probably costs, against what you can actually price

Since Alibaba will not give a number, use the model it succeeds. Qwen3.7 Max lists at $2.50 input and $7.50 output per million, routinely discounted to $1.25/$3.75 on OpenRouter. A Max-tier successor rarely lands cheaper on its published rate, so treat $2.50/$7.50 as a floor for whatever Qwen3.8 Max eventually prints, not a promise. Here is where that anchor sits on a 20M-input, 5M-output month, next to the flagships that already have rate cards.

ModelInput / 1MOutput / 1M20M / 5M month
Claude Fable 5$10.00$50.00$450
GPT-5.6 Sol$5.00$30.00$250
Kimi K3$3.00$15.00$135
Qwen3.7 Max (the anchor)$2.50$7.50$87.50
GLM-5.2$1.40$4.40$50
DeepSeek V4-Pro$0.435$0.87$13.05
Qwen3.8 Max Previewunpricedunpricedcredits only

The point of the table is the bottom row. Every model above it can be dropped into a calculator and turned into a monthly bill. Qwen3.8 Max cannot, because it does not compete on this axis yet. If the successor holds the 3.7 Max line, it would run around $87 to $90 for that month, which slots it between Kimi K3 and GLM-5.2 and makes it a mid-priced frontier option rather than a cheap one. But that is a projection off last quarter's model, not a rate you can plan against.

Nor does the credit plan close the gap. Even the ~2% overnight math only tells you the rate is aggressive, not what a unit of work costs, because the credit-to-token bridge is missing. You could burn a Pro plan's 40,000 weekly credits fast or slow and have no way to model it in advance. For a team that needs a forecastable line item, that uncertainty is itself a cost.

Why Alibaba priced it this way

The timing explains the packaging. Qwen3.8 Max landed two days after Moonshot released Kimi K3 as open weights, a 2.8T model that reset expectations for how good a downloadable Chinese frontier model can be. Answering that with a preview, a subscription, and a headline claim lets Alibaba plant a flag at WAIC without committing to a number it would have to defend on a public leaderboard or a public rate card. A credit plan with a night discount also nudges usage onto spare overnight capacity, which is a sensible way to stress-test a giant model before you promise anyone an SLA.

None of that is nefarious, but it does shift risk onto the buyer. The open-weight wing of the Chinese market, DeepSeek, GLM, Kimi, competes on transparent per-token rates and, increasingly, downloadable weights. Qwen3.8 Max competes on a promise, priced in a unit you cannot convert. That is a fine way to launch a preview and a poor way to win a line in a cost model.

Should you wait for it

If you are exploring and the credit plan fits your budget, the overnight discount makes Qwen3.8 Max cheap to poke at, and the new multimodal input is a real reason to try it on image and document work the text-only 3.7 Max could not touch. Treat that as an experiment, not a migration. You are testing capability, not locking in a cost.

If you are choosing a production model against a budget, you cannot pick this one yet, and that is the correct call rather than a cautious one. There is no rate to compare, no benchmark to justify a premium, and no open weight to self-host. Everything you would need to defend the choice in a spend review is missing. The models directly around its likely price, Kimi K3 at $3/$15 and GLM-5.2 at $1.40/$4.40, publish both their rates and their scores, so they can be reasoned about today.

Watch for two things before you take Qwen3.8 Max seriously as a budget line: a per-million-token rate on Model Studio, and an independent Intelligence Index score. When both exist, drop them into the calculator against your real token mix and see where it actually lands. Until then, the pricing page has every model that will give you a number, and the Qwen3.7 Max write-up covers the model this one is built on. For the open-weight rival it is answering, see the Kimi K3 pricing piece.

Sources