Skip to main content
TokenCost logoTokenCost
ComparisonOctober 3, 2026·6 min read

Alibaba charges double for Prime on Qwen3.8 Max and GLM-5.3. On Qwen it buys 1.6x to 1.9x the speed. On GLM-5.3, Alibaba's own standard endpoint is faster, and Decart is 3.4x faster for 42% of the price.

Same tier name, same 2x price, same 1.5x-to-2x promise from Alibaba. Ten days of traffic later, one of them delivers it and the other is slower than the thing it replaces.

Red and cyan light trails streaking down a dark highway at night, long-exposure motion blur

Photo by Jahanzeb Ahsan on Unsplash

If you run Qwen3.8 Max and your users are waiting on it, buy Prime. You pay $4 and $12 instead of $2 and $6, and you get somewhere between 58 and 70 tokens a second instead of 36, with the first token arriving in under a second instead of four. There is no other host for Qwen3.8 Max, so this is the only speed you can buy.

If you run GLM-5.3, don't. Its Prime endpoint measured 64 tokens a second, below Alibaba's own standard GLM-5.3 endpoint at 75. The weights are open, 39 endpoints serve them on OpenRouter, and the fastest one does 218 tokens a second at well under half the Prime price.

What Prime is, and who sells it

Alibaba Cloud announced Prime mode at its Yunqi conference on September 22. It is not a new model. It is the same weights on faster serving, and Alibaba's Prime mode page promises throughput of 1.5 to 2 times the standard API. Both Prime models reached OpenRouter on September 23.

One thing tripped us up. GLM-5.3 Prime sounds like something Z.ai sells. It isn't. Z.ai's price list has no Prime, Fast or Turbo line. GLM-5.3 Prime is Alibaba Model Studio serving Z.ai's open weights, and Alibaba is the only provider behind it on OpenRouter.

Per 1M tokensInputCachedOutput
Qwen3.8 Max Prime$4.00$0.50$12.00
Qwen3.8 Max$2.00$0.25$6.00
GLM-5.3 Prime$2.80$0.56$8.80
GLM-5.3 (Z.ai list)$1.40$0.26$4.40
GLM-5.3 (Alibaba standard)$1.19$0.238$3.74

OpenRouter model and endpoint listings, read October 3, 2026; Z.ai pricing page for the list row. All five have a 1M-token window and 131,072 max output.

On Alibaba's own price list both Primes are exactly 2x their standard sibling in every column, in yuan. On OpenRouter the GLM numbers drift: Alibaba sells standard GLM-5.3 there at $1.19, a bit under Z.ai's $1.40, but prices Prime at twice Z.ai's card. Measured against Alibaba's own standard endpoint, GLM-5.3 Prime costs 2.35x.

Qwen: 2x the bill, 1.6x to 1.9x the speed

OpenRouter publishes median throughput and time to first token for every endpoint. On the morning of October 3 it had 329 Prime requests and 9,869 standard ones for Qwen3.8 Max.

Qwen3.8 MaxTokens/s (p50)First token (p50)First token (p90)
Prime70962 ms2,451 ms
Standard363,914 ms19,047 ms

OpenRouter endpoint stats, read 06:06 UTC October 3, 2026. OpenRouter does not state the window these medians cover.

Seven minutes later we pulled the page again. Prime had nearly doubled its sample to 618 requests and its median had dropped to 58 tokens a second, with first token at 885 ms; standard hadn't moved. So call the throughput gain 1.6x to 1.9x, inside Alibaba's range, on a sample that is still small and still settling. The bigger change is the wait. Standard Qwen3.8 Max takes almost four seconds to start talking on a typical request and 19 on a bad one. Prime is under a second and two and a half.

Put a number on it. A request with 20,000 tokens in and 2,000 out costs $0.052 on standard and finishes in about 59.5 seconds. On Prime it costs $0.104 and finishes in 29.5 to 35.4 seconds, depending on which reading you trust. You pay 5.2 cents to get 24 to 30 seconds back. For a chat window or a coding agent someone is watching, that is cheap. For an overnight batch it is money for nothing.

GLM-5.3: Prime is the slow lane

Here the comparison is not Prime against one standard endpoint. It is Prime against everyone else who serves the same weights. Below is the same 20,000-in, 2,000-out request run 1,000 times, plotted as seconds per request with the bill for the thousand alongside.

  • Decart (FP4)9.6s · $31.28
  • Mistral15.3s · $36.80
  • Baseten fast (FP8)17.4s · $55.20
  • Alibaba standard28.1s · $31.28
  • Alibaba Prime32.2s · $73.60
  • Z.ai first-party43.4s · $36.80

Seconds = median time to first token + 2,000 / median tokens per second, from OpenRouter endpoint stats read October 3, 2026. Dollars are list input and output prices, no caching. Medians: Decart 218 tok/s, Mistral 142, Baseten fast 119, Alibaba standard 75, Prime 64, Z.ai 49.

Prime is the most expensive row on the chart and the second slowest. Alibaba's own standard endpoint beats it on both. The only host it outruns is Z.ai, the lab that made the model, which runs its first-party API at 49 tokens a second.

Decart is the outlier. It charges exactly what Alibaba's standard endpoint charges, $1.19 and $3.74, and runs 3.4 times as fast as Prime. The thousand requests cost $31.28 against Prime's $73.60. Mistral, at Z.ai's list price, is more than twice as fast as Prime too.

There is a catch, and it's the reason we'd test before switching. Decart serves FP4 weights and Mistral serves NVFP4. Lower precision is a big part of how they get the speed, and it can cost accuracy on hard reasoning. If you want a quick host without that question, Baseten's FP8 fast endpoint does 119 tokens a second at $2.10 and $6.60, still nearly twice Prime's speed for 25% less money.

How this lines up with OpenAI's speed tiers

Pay-for-speed is everywhere now. OpenAI's Fast tier is 2x, Anthropic's fast mode for Opus is about 2x, and xAI's Grok 4.7 Fast is 2x, though only inside Cursor and Grok Build. At DevDay on September 29 OpenAI added Ultrafast at 6x, for GPT-6 Astra only: $60 in, $6 cached and $300 out per million tokens. OpenAI's launch post says up to 6x faster in the API.

The same OpenRouter stats let us check Astra. Its Fast endpoint ran at a median 50 tokens a second against 43 on standard, from 250 requests. That's 1.16x the speed for 2x the price, a long way short of the "up to 2.5x" OpenAI has advertised for Fast. Ultrafast had served 22 requests through OpenRouter by our second read, at a median 33.5 tokens a second, slower than standard. That is too few requests to judge, and we'd expect it to settle well above that.

So by the only independent numbers we have, Qwen Prime is the best deal of these speed tiers, which isn't what we expected to write. It's also the one with the smallest sample. We'll rerun this once Prime has a few thousand requests behind it.

What the price list leaves out

Alibaba's own price list shows Qwen3.8 Max Prime in the Beijing region only, at CNY 24 and CNY 72 per million, which it converts to $3.301 and $9.902. The $4 and $12 on OpenRouter are twice the Singapore standard rate. If you call Alibaba directly, the endpoint is a cn-beijing host.

Prime does not change the window or output limit: still 1M tokens in and 131,072 out on both models. Neither Prime has an Artificial Analysis page yet, so there are no independent quality scores for them; the weights are the same, so they should match standard Qwen3.8 Max and GLM-5.3, which both sit at 45 on the Intelligence Index.

One oddity: Alibaba's Prime mode page lists glm-5.3-prime but no Qwen model. qwen3.8-max-prime appears only on the pricing page, which links to Prime mode. Both work on OpenRouter today.

One lever against thirty-nine

A speed tier is worth paying for when there is no other way to get the speed. Qwen3.8 Max has closed weights and one host, so Prime is the only lever, and it works. GLM-5.3 has open weights and dozens of hosts, so Prime is competing with the open market and losing.

Compare both models and their hosts in the calculator, or see the base rates on the Qwen3.8 Max and GLM-5.3 pages.

Where the numbers come from