Skip to main content
TokenCost logoTokenCost
Model ReleaseAugust 4, 2026·11 min read

Qwen3.8-Max prices output 60% under Kimi K3. On the same benchmark suite the bill came in 11% lower, and the $2 sticker is not on any page Alibaba owns.

Alibaba shipped its 2.4 trillion parameter flagship to general availability yesterday and the number everyone printed is $2 in, $6 out per million. It is a real price. You can spend money at it this morning on eight different gateways. It is just not on Alibaba's pricing page, which still tops out at Qwen3.7-Max and currently runs that model at an effective $1.25 and $3.75, undercutting its own successor. Meanwhile the resellers quoting different figures turn out not to disagree at all: AIHubMix sits at 0.845x the baseline on both legs, Deep Infra at 0.825x on both. One rate card, wholesale discounts on top. The number that actually moves is the one nobody put in a headline. Artificial Analysis clocked this model generating 150M tokens to complete its Intelligence Index against a 63M median, so a card priced 60% below Kimi K3 on output delivered a bill 11.4% smaller rather than 60%. Push that correction through an ordinary agent workload and the distance to GPT-5.6 Sol collapses from more than triple to a difference of 44%.

Amber light trails woven into a dense grid on black, suggesting sparse routing through a large network

Photo by Puru Raj on Unsplash

Three claims, and how much weight each one holds

This launch has an unusual amount of secondhand information attached to it, so before anything else, here is where each of the load-bearing facts in this post comes from and how far I would trust it.

  • $2 and $6 is what you will pay, but Alibaba has not said so. Read straight off the OpenRouter models API and matched by seven other gateways. Solid enough to budget against. Just be aware there is no first-party rate card behind it, and the predecessor is currently cheaper.
  • The verbosity gap is measured, not estimated. 150M tokens to complete the Intelligence Index against a 63M median, and a $2,159.51 bill for doing it. Artificial Analysis ran it, published both figures, and the comparison to Kimi K3 falls out of their numbers rather than mine.
  • The 80% overnight discount everyone is quoting has expired. Confirmed against the promotion page itself. At general availability it is 50%, and it only ever applied to credit subscriptions rather than to per-token billing.

The price is real. Alibaba just has not published it.

I went looking for the first-party rate card and could not find one. Model Studio's pricing page lists Qwen3.7-Max as the newest Max variant. The English models page does not mention 3.8 either. The Chinese docs site does list the model IDs and the regional endpoints, which is how you know the thing exists and where to call it, but there is no price attached. The context-cache documentation, which enumerates every model eligible for caching, leaves 3.8 out of the list entirely.

To be fair to the number, Artificial Analysis describes its $2 and $6 as based on Alibaba's API, and models.dev carries two Alibaba first-party rows for the model. Both of those show $0.00 per token, because what Alibaba sells directly at the moment is a credit subscription rather than metered tokens. So the rate almost certainly originates with Alibaba. It just has not been written down anywhere a customer can point at.

This matters less than it sounds and more than you would expect. Less, because $2 and $6 is what you will actually be charged, on OpenRouter and Vercel and six other gateways, and their numbers agree to the cent. More, because the predecessor is currently cheaper. Qwen3.7-Max lists at $2.50 and $7.50 with a limited-time 50% cut applied, which puts it at $1.25 and $3.75. Buying the new flagship therefore costs 60% more than buying the model it replaces, and the only reason that comparison is available at all is that Alibaba published one of the two.

Where you buy itInputOutputMultiple of baselineContext sold
OpenRouter, Vercel, Kilo, Merge, EmpirioLabs$2.00$6.001.000x1M
NanoGPT$2.00$6.001.000x991K
CrossModel$1.88$5.630.94x1M
AIHubMix$1.69$5.070.845x991K
Deep Infra$1.65$4.950.825x256K only
Qubrid, after 20% launch discount$2.30$5.691.15x / 0.95x1M
Alibaba Model Studionot publishednot published-1M

Read the fourth column before you conclude that anyone is competing on price. AIHubMix is exactly 0.845 times the baseline on input and exactly 0.845 on output. Deep Infra is 0.825 on both. CrossModel is 0.94 on both. Those are not independent pricing decisions, they are flat wholesale margins applied to one number, which is itself decent evidence that $2 and $6 is Alibaba's real rate even though Alibaba has not said so. Only Qubrid built its own card, marking input up 1.435x and output only 1.187x before discounting, and Deep Infra is the one row where the cheaper price buys a different product: 256K of context instead of a million.

It writes 150 million tokens to answer the same questions

Artificial Analysis runs every model it covers through the same nine-eval Intelligence Index, and it reports two things people mostly ignore: how many tokens the model burned getting through the suite, and what that cost. For Qwen3.8-Max the answer is 150M tokens and $2,159.51, against a median of 63M. Their summary calls it "very verbose" and "reasonably priced" in the same breath, which is a more useful review than most write-ups managed.

ModelTokens to finish the indexAgainst the 63M medianOutput rate
Motif 3 (Beta)200M3.17x-
Qwen3.8-Max150M2.38x$6
Kimi K3 (max)130M2.06x$15
Claude Opus 4.8 (max)120M1.90x$25
GPT-5.6 Terra (max)96M1.52x$12
GPT-5.6 Sol (high)21M0.33x$30
GPT-5.6 Sol (low)6.6M0.10x$30

One methodological note, because it would be easy to misuse this table. Artificial Analysis quotes a median relative to whatever peer set a given model page is drawn against, and the Kimi K3 page cites a 100M median rather than 63M. I have used 63M as the single denominator throughout, since that is the figure carried on the GPT-5.6, Opus 4.8 and Motif 3 pages, and a ratio is only meaningful if every row divides by the same thing. The raw token counts are direct measurements and are not affected.

A 60% discount that arrived as 11.4%

Kimi K3 is the natural comparison. Both are trillion-parameter sparse MoE flagships with a 1M window and no long-context surcharge, released about two and a half weeks apart, and both come out of China. On the rate card Qwen wins by a street: $2 against $3 on input, $6 against $15 on output. That output line is 60% cheaper.

Then you look at what the two of them actually cost to run through an identical workload, which is the one experiment nobody but a benchmark lab bothers to do.

MeasureQwen3.8-MaxKimi K3Qwen advantage
Output rate per 1M$6.00$15.0060.0% cheaper
Blended 3:1$3.00$6.0050.0% cheaper
Bill for the full index$2,159.51$2,437.4111.4% cheaper
Intelligence Index53574 points behind
Cost per index point$40.75$42.764.7% cheaper
Output speed46.5 tok/s~36 tok/s29% faster

Sixty percent off the output line, eleven percent off the invoice. The gap between those two numbers is the whole reason this post exists. Divide each bill by the tokens it produced and Qwen3.8-Max lands at $14.40 per million generated against Kimi's $18.75, so even after normalising for volume the advantage is 23% rather than 60%, because the verbose model also drags more input back through the meter on every retry and continuation. Cost per index point is the least flattering cut of all: $40.75 against $42.76, a difference of under 5% for four points of measured intelligence.

Speed is the one row that goes the other way, and it goes there hard. Both models are slow against a 63 tokens per second median, but Kimi K3 runs at roughly 36 and Qwen3.8-Max at 46.5, so the cheaper model is also about 29% faster. If you came here expecting the usual tradeoff where the budget option costs you latency, this is not that.

Repricing a normal job by measured token volume

Benchmark suites are not your workload, so here is the same correction applied to something ordinary. Take an agent job that pulls 20M tokens of input over a month and would draw 4M tokens of output from a model of median verbosity. Price it twice: once the way everybody prices things, multiplying list rates by identical token counts, and once with each model's output scaled by the volume Artificial Analysis actually measured.

ModelRateAssuming 4M outOutput it would writeAdjusted bill
Qwen3.8-Max$2 / $6$64.009.52M$97.14
GPT-5.6 Terra$2 / $12$88.006.10M$113.14
GPT-5.6 Sol, high effort$5 / $30$220.001.33M$140.00
Kimi K3$3 / $15$120.008.25M$183.81
Claude Opus 4.8$5 / $25$200.007.62M$290.48

Sol goes from the most expensive row on the table to the middle of it. Qwen3.8-Max stays cheapest, which is worth saying plainly, but its margin over Sol goes from more than triple to a difference of 44%, and Terra ends up within sixteen dollars of it. That is a different shopping decision than the one the rate card implies, and it is entirely driven by a column most pricing pages do not have.

Treat the adjusted column as a directional estimate rather than a quote. Verbosity measured on a reasoning benchmark will not transfer perfectly to your prompts, and Sol's low-effort setting writes a tenth of the median, so the effort dial moves this more than the model choice does. The point is not that these are your numbers. It is that four times out of five, the unadjusted column is the only one anyone runs, and it is wrong in a consistent direction: it flatters verbose models with cheap output rates. You can run your own version of this on the cost calculator by putting your real output counts in rather than a guessed ratio.

Two different night discounts are being quoted as one

Almost every writeup of this launch mentions an 80% overnight discount, and the reason the claim is so slippery is that there are two unrelated mechanisms with similar numbers attached, and coverage keeps welding them together.

The first is real, first-party and genuinely per-token. Alibaba attaches a limited-time night discount of 80% off, between 22:00 and 08:00 UTC+8, to Qwen3.7-Max on the Hong Kong, Frankfurt and US-Virginia endpoints. That is a token price, not a credit adjustment, and it is on Alibaba's own pricing table. The catch is the model name: it is documented for 3.7, and since Alibaba has published nothing at all for 3.8, there is no way to confirm it carries over.

The second is a credits promotion on Qoder, and this is the one people are actually citing. At general availability the multiplier is 0.5x in regular hours dropping to 0.25x between 22:00 and 08:00 Singapore time, so 50% off overnight, and the page says plainly that discounts affect credits pricing only. During the preview it was far steeper, roughly 90% all day with the off-peak stacking on top to reach about 98%, which is where the eye-catching figures in circulation come from. That preview ended on August 3.

So if you are buying credits through Qoder, the overnight saving today is half, not four fifths. If you are buying tokens, the 80% is documented for the previous model rather than this one, and you should verify it against your own endpoint before planning around it. Compare that to DeepSeek putting time-of-day pricing on its API, where one published table covers one billing mode and there is nothing to disentangle.

The cache rates break Alibaba's own documented percentages

Alibaba's context-cache documentation is specific: explicit cache creation bills at 125% of the standard input rate, an explicit cache hit at 10%, an implicit hit at 20%. Apply those to a $2 input price and you should see $2.50 to create, $0.20 on an explicit hit, $0.40 on an implicit one.

What the gateways carry is $2.50 to create and $0.25 to read. The creation fee matches perfectly. The read does not: $0.25 is 12.5% of $2, which is neither of the two documented percentages. AIHubMix quotes $0.17 and Qubrid $0.16, and at least one API aggregator states $0.20 outright, contradicting the rest.

There is an arithmetic coincidence here that I want to flag as a hypothesis rather than a finding, because it would explain everything and I cannot confirm it. $0.25 is exactly 10% of $2.50, and $2.50 is exactly what Qwen3.7-Max lists at before its limited-time 50% cut. If Qwen3.8-Max's true list price is also $2.50 and the $2 circulating everywhere is itself a promotional rate, then the cache numbers snap into place against the documented 10%. That would also mean the price rises 25% whenever the promotion ends. I have no evidence for it beyond the arithmetic, and Alibaba publishing an actual rate card would settle it in a sentence.

Every benchmark on the launch page is vendor-run

Alibaba published a proper evaluation table this time, which is a real improvement on the July preview, where the company claimed second place in the world behind Claude Fable 5 and supplied no scores, no parameter count and no price to support it. The numbers are still all its own.

BenchmarkQwen3.8-MaxClaude Fable 5Claude Opus 4.8GPT-5.6 Sol
Terminal-Bench 2.186.684.684.688.8
SWE-bench Pro67.780.069.264.6
PaperBench93.088.880.390.5
IFBench82.863.562.272.7
GPQA Diamond92.692.692.094.1
HLE43.653.345.747.2
OSWorld-Verified86.1not publishednot publishednot published

Worth noting how many rows Alibaba did not lead on its own launch table. Four of the six: Terminal-Bench, GPQA Diamond and HLE all go to someone else, and SWE-bench Pro is the one that should give a buyer pause, because it is the benchmark most closely tied to the agentic coding work this model is being sold for. 67.7 puts it behind Opus 4.8 and a long way behind Fable 5. Set against that, IFBench is a genuine standout: 82.8 beats the next best model on the table by ten points and the weakest by twenty. Instruction following looks like the real strength here and is worth buying for. Repo surgery, on the vendor's own evidence, is not.

The independent picture is thinner and roughly agrees. Artificial Analysis puts Qwen3.8-Max at 53 on Intelligence Index v4.1, against 57 for Kimi K3, 59 for GPT-5.6 Sol and 61 for Claude Opus 5. Arena WebDev has 1,668 against Kimi K3's 1,676, close enough that the error bars overlap. And when Eden AI tried to line the two vendors' own tables up, they reported finding only a handful of benchmarks where the two harnesses actually matched, TerminalBench 2.1 among them. I have not been able to reproduce that audit independently, but the warning it points at is sound. If you see a fifteen-row comparison table declaring a winner this week, check whether the rows came from one harness or two.

Specs, and the four things still missing

2.4T total parameters with 95B active, so about 4% of the pool fires on any token. Text, image and video in, text out. A reasoning budget of 262K tokens, a reasoning_effort control with low, medium and xhigh settings and xhigh as the default, and thinking tokens billed as output at $6. Endpoints in Beijing, Singapore, Tokyo, Frankfurt and US-Virginia, with OpenAI-, Anthropic- and DashScope-compatible APIs. Live on OpenRouter and Vercel AI Gateway; not on Bedrock or Vertex.

The context window is quoted four different ways depending on where you look, and the spread is mostly not an error. OpenRouter and models.dev both report 1,000,000 with 131,072 max output, and OpenRouter is the primary source here since it is an API response rather than a marketing page. AIHubMix says 991K, MarkTechPost says 991K usable input dropping to 983K with thinking on. Those describe the same thing: what is left for input after output and reasoning are reserved out of the same pool. Vercel's 128,000 max output and one aggregator's 65,536 are the two figures I would ignore.

Three things I went looking for and did not find. There is no published knowledge cutoff and no model card, which is unusual for a GA flagship. Batch pricing is documented at 50% off for Qwen3.7-Max but unconfirmed for 3.8. And the weights, promised on Hugging Face and ModelScope alongside a Qwen3.8-27B, had not appeared as of this morning, so the model is closed for now and no third party can evaluate it independently. Rate limits are the one gap that has been filled: 2M tokens per minute and 15,000 requests per minute, though not on an Alibaba page.

Buy it for input, not for output

The shape of the decision is unusually clean. At $2 input with a flat million-token window and no tiering, this is one of the cheapest ways to push a lot of context through a frontier-class model, and the 92.6 GPQA and 82.8 IFBench say it will do something useful with it. Document extraction, long-context retrieval, classification over big corpora, anything where the prompt is enormous and the answer is short: the verbosity problem barely touches you, and you should try it this week.

Invert the ratio and the case weakens fast. An agent loop that reasons at xhigh by default, on a model that writes 2.38x the median, is the exact workload where a $6 output rate stops behaving like a $6 output rate. Turn the effort dial down before you benchmark it, because that single setting moved the numbers in my table further than picking a different vendor did.

Whatever you do, log output tokens per completed task rather than per request, and compare that against your current model before you migrate anything. It is the only number that would have caught the 60%-becomes-11% gap in advance, and almost nobody tracks it. You can put the results side by side on our model comparison pages once you have them. I would also wait a week before committing anything large: the weights are supposed to land shortly, and a first-party rate card would resolve both the cache arithmetic and the question of whether $2 is the price or the promotion.

Sources

  • OpenRouter: models API - The primary source for $2.00 input, $6.00 output, $0.25 cache read and $2.50 cache write, plus 1,000,000 context with 131,072 max completion tokens, the August 3 listing date, and the text, image and video input modalities
  • Alibaba Cloud Model Studio: model pricing - Checked August 4 and carrying no entry for Qwen3.8-Max. Qwen3.7-Max is the newest Max listed, at $2.50 and $7.50 with a limited-time 50% discount, no context-length tiering, a 50% batch discount and 1M free tokens valid 90 days
  • Artificial Analysis: Qwen3.8-Max - Intelligence Index 53, 150M tokens generated across the index against a 63M median, $2,159.51 to run the suite, 46.5 tokens per second output and 2.48s to first token, plus the "very verbose" and "reasonably priced" assessments
  • Artificial Analysis: Kimi K3 - Index 57, 130M tokens across the index, $2,437.41 to run it, roughly 36 tokens per second against a 63 median, and the 100M peer median that differs from the 63M used elsewhere. Several secondary write-ups quote $2,690.80 and 62 tokens per second for this model; both figures disagree with the page above, which is what the numbers in this post use
  • Qoder: Qwen Max GA promotion - The actual off-peak terms: a 0.5x credits multiplier in regular hours dropping to 0.25x between 22:00 and 08:00 Singapore time, running from August 3, applying to credit plans only, with the explicit note that discounts affect credits pricing rather than the model
  • Alibaba Cloud: context cache - The documented 125% creation, 10% explicit-hit and 20% implicit-hit percentages that the circulating $0.25 cache read fails to match, and a supported-model list that does not include Qwen3.8-Max
  • models.dev: Qwen3.8-Max providers - The nine-provider spread showing CrossModel at $1.88/$5.63, Deep Infra at $1.65/$4.95 capped to 256K context, and five gateways at exactly $2/$6, which is what makes the flat-margin reading obvious
  • AIHubMix: Qwen3.8-Max - $1.69 and $5.07 with a $0.17 cache read and 991,000 context, which works out to exactly 0.845x the baseline on both the input and output lines
  • Qubrid AI: Qwen3.8-Max API - The only genuinely independent rate card at $2.87 and $7.12 list, 20% off at launch, plus confirmation that pricing is flat across the full 1M context with no multiplier past 200K, and the observation that every published benchmark is vendor-run
  • MarkTechPost: Alibaba releases Qwen3.8-Max - The 2.4T total and 95B active parameter counts, the 262K reasoning budget, the reasoning_effort settings, the 991K and 983K usable input figures, the regional endpoint list, and Alibaba's published benchmark table