Two models called Flash arrived five and a half hours apart yesterday. One of them looks 2.13 times cheaper, and on the only workload anybody has measured they bill thirty cents apart.
Z.ai listed GLM-5.3-Flash at 13:59 UTC and Alibaba listed Qwen3.8-Flash at 19:37. Every tracker now shows the first at $0.075 per million input tokens and the second at $0.16, which reads as a rout. It is a fortnight-long one. The $0.075 is half of a list price that both Z.ai's own pricing page and a single field in OpenRouter's endpoint API will tell you about, and the promotion behind it expires on September 9. Price the same measured run on the card that replaces it and the two models finish $0.30 apart on $138.

Photo by Jerry Wei on Unsplash
The same run, on both sides of one date
Today, through September 9
$69.01
GLM-5.3-Flash on the promotional card, against Qwen3.8-Flash's $137.72. Very close to two-for-one, and the reason every comparison published this week reads as a landslide.
From September 10
$138.02
The same tokens on GLM-5.3-Flash's list card, against the same unchanged $137.72. Qwen3.8-Flash is now the cheaper of the two by $0.30, and by more than that if you cache.
Both columns price an identical 420.13 million input and 150 million output tokens, which is not a workload we invented: it is what Artificial Analysis actually spent scoring GLM-5.3-Flash, recovered from its own published total. The derivation is below.
Five hours and thirty-eight minutes apart
The collision is real rather than a framing device. OpenRouter stamps every listing with a creation timestamp, and the two decode to August 26 at 13:59 and 19:37 UTC. Two labs put a cheap multimodal reasoning model with a million-token context on sale on the same afternoon, and the resemblance runs deeper than the launch window: both take text, image and video in, both cap a single response at 131,072 tokens, both are aimed squarely at long-horizon agent work.
| Field | GLM-5.3-Flash | Qwen3.8-Flash |
|---|---|---|
| Listed on OpenRouter | Aug 26, 13:59 UTC | Aug 26, 19:37 UTC |
| List price, in / out | $0.15 / $0.50 | $0.16 / $0.47 |
| Price today, in / out | $0.075 / $0.25 | $0.16 / $0.47 |
| Cached input | $0.03 list, $0.015 now | $0.016 |
| Cache write | Not published | $0.20 |
| Context window | 1,048,576 | 1,000,000 |
| Max output | 131,072 | 131,072 |
| Weights | MIT, 320B-A18B | None on this listing, 125B MoE |
| Reasoning | Mandatory, defaults to max | Optional, budget-cappable |
| Implicit caching | False on all 9 sellers | True |
| Sellers | 9 | 1 (Alibaba) |
Three rows in that table decide more money than the price rows do, and we will come back to each: who can turn reasoning off, whose cache is automatic, and how many companies you can buy the thing from. Start with the price rows anyway, because one of them is not what it appears to be.
The cheap number has a flag attached to it
OpenRouter's per-model endpoint API returns a discount field alongside each seller's rates. On Z.ai's own endpoint for GLM-5.3-Flash it reads 0.5. That single number says the $0.075 you see quoted everywhere is not the price of the model, it is half the price of the model, and there is a list rate sitting behind it waiting to come back.
You do not have to take the field's word for it. Divide the quoted rate by one minus the discount and you get $0.15 input and $0.50 output, and Z.ai's own pricing table prints exactly those two figures as the list price with the promotional pair beside them, and dates the promotion's end at September 9, 2026, 24:00 UTC+8. Two independent sources, the same two numbers. The arithmetic on the discount field is sound, which matters for the next section because five other sellers let us run it in reverse.
This is the second time in four days that the discount field has settled something a model card would not. On August 24 it read 0 for the stealth model then called Ox Alpha, which is how we knew that model's $0 was a genuine list price rather than a markdown. Here it reads 0.5, and the conclusion runs the other way.
Line by line, the cheaper model changes twice
Set the promotion aside and compare the cards that will be in force in a fortnight. Read down the rows rather than across the models, because the interesting thing about these two rate sheets is that neither one wins.
| Rate card line | GLM now | GLM from Sep 10 | Qwen | Cheaper at list |
|---|---|---|---|---|
| Input | $0.075 | $0.15 | $0.16 | GLM, 1.067x |
| Cached input | $0.015 | $0.03 | $0.016 | Qwen, 1.875x |
| Output | $0.25 | $0.50 | $0.47 | Qwen, 1.064x |
| Cache write | n/a | n/a | $0.20 | Not comparable |
GLM-5.3-Flash keeps a 6.7% edge on fresh input and gives up 6.4% on output, and output is the bigger line on any reasoning model. On cached input it is beaten by a wide margin, because the two labs chose different discounts off their own headline rate: Alibaba charges a tenth of its input price for a cache read, Z.ai charges a fifth. That 10x against 5x is the largest structural difference between the two cards and the only one that survives the promotion in either direction.
Artificial Analysis published a bill. Run it backwards and you get a token count.
Comparing rate cards on a workload somebody invented is the weakest move in this genre, and we have made it often enough. There is a better option here, because one real workload has already been run on one of these two models and costed in public. Artificial Analysis reports that scoring GLM-5.3-Flash on its Intelligence Index consumed 150 million output tokens and cost $138.02, and that it billed the run at the $0.15 and $0.50 list card rather than the promotion.
Those three facts are enough to recover the fourth. Output at $0.50 accounts for $75.00 of the total, so the remaining $63.02 was input, and $63.02 at $0.15 per million is 420.13 million input tokens. The suite ran at 2.80 input tokens for every output token. Now there is a measured workload to price, and the only assumption left is that Qwen3.8-Flash would spend the same tokens on the same questions, which we flag as an assumption and revisit two sections down.
420.13M input and 150M output, billed four ways:
- GLM-5.3-Flash, promo
- $69.01
- Qwen3.8-Flash
- $137.72
- GLM-5.3-Flash, list
- $138.02
- GLM-5.3 flagship
- $1,248.19
The two highlighted bars are the whole argument. They differ by about $0.30 on $138, which is 0.22%, and no reader could pick the cheaper one off a chart. Artificial Analysis publishes its inputs rounded, so treat that gap as roughly thirty cents rather than exactly thirty.
Rounding is worth a word here, because the gap is small enough to care about it. Artificial Analysis publishes 150M and $138.02, both rounded, so the 420.13M we recovered is good to about half a percent, and that band is wider than the thirty cents itself. The direction survives anyway, and there is a cleaner way to see why. GLM-5.3-Flash is $0.01 cheaper per million on input and $0.03 dearer on output, so the two cards tie at exactly three input tokens per output token. Above that ratio GLM wins, below it Qwen does, and the measured suite ran at 2.80. That is the finding stated in a form no rounding can move.
The bottom bar is worth a sentence of its own. Ten days before its Flash sibling appeared, GLM-5.3 was billing $1.40 and $4.40, and the same tokens on that card come to $1,248.19. Z.ai has cut the price of its own family by 9.04x in a week and a half, for a model Artificial Analysis scores at 57 against the flagship's 60. Whatever else is true about the Qwen comparison, that is the more consequential number on this page for anyone currently paying the flagship rate.
Flash is a claim about the price, not about the speed
Both of these models are named after a quality neither of them has relative to the tier above. Artificial Analysis clocks GLM-5.3-Flash at 50.2 output tokens per second and GLM-5.3 at 85.3, so the flagship is 1.70x faster than the model whose name means quick. Time to first token is near enough identical at 1.47 seconds against 1.57.
Put that beside the token counts and the trade gets sharper. The Flash model emits 150 million output tokens across the index and the flagship emits 170 million, a difference of only 11.8%, so the cheaper model is barely more concise. Divide each by its own throughput and the index takes roughly 830 hours of generation on GLM-5.3-Flash against 554 on GLM-5.3. You are buying an 8.97x smaller bill and paying about half as long again in wall-clock for it. On a nightly batch job that is free money. On anything a person waits for, it is not.
Qwen3.8-Flash has no Artificial Analysis entry yet, so there is no like-for-like figure to set against those. OpenRouter's own telemetry puts it at 52 tokens per second with a median latency of 4.68 seconds, which is a different measurement taken a different way and should not be put in the same column. The throughput figures happen to land close together. The latency figures are not comparable at all, and we are not going to pretend otherwise.
Only one of them lets you turn the meter down
Here is the row that outweighs the rate cards. GLM-5.3-Flash's listing sets mandatory to true on its reasoning block and default_effort to max, with supported efforts of max, high and low. There is no off. Qwen3.8-Flash reports mandatory false and advertises support for a reasoning token budget, so you can put a ceiling on the thinking directly.
Reasoning tokens bill as output tokens, and output is 54.3% of GLM-5.3-Flash's measured index bill. So the question of who is cheaper stops being a question about $0.50 against $0.47 the moment one party can halve the quantity and the other cannot.
| If Qwen emits | Output tokens | Qwen bill | vs GLM at list |
|---|---|---|---|
| The same as GLM | 150M | $137.72 | 1.002x cheaper |
| Three quarters | 112.5M | $120.10 | 1.149x cheaper |
| Half | 75M | $102.47 | 1.347x cheaper |
| A quarter | 37.5M | $84.85 | 1.627x cheaper |
None of those rows is a measurement, and quality would move with the budget. That is the point rather than a weakness in it: the rows describe a lever one vendor gives you and the other does not, and the lever's range is wider than the entire gap between the two rate cards. GLM-5.3-Flash's equivalent move is to drop from max effort to low, which is real and worth doing, but its floor is a reduced amount of thinking rather than none.
Ten sellers, one set of weights, a clean factor of two
The MIT license did in a day what it took GLM-5.2 weeks to accumulate. We counted ten sellers this morning and nine this afternoon, which is its own kind of answer. That is the structural advantage over Qwen3.8-Flash, which is closed and sold by Alibaba alone. MarkTechPost's launch write-up has the architecture detail we are not going to duplicate here: 320 billion parameters total, 18 billion active. It is also where the promotion gets confusing, because only three of the nine are running it.
| Seller | Rate today, in / out | Discount | The measured run |
|---|---|---|---|
| Z.AI | $0.075 / $0.25 | 50% | $69.01 |
| Novita | $0.075 / $0.25 | 50% | $69.01 |
| GMICloud | $0.075 / $0.25 | 50% | $69.01 |
| Venice | $0.09375 / $0.3125 | None | $86.26 |
| Parasail | $0.15 / $0.50 | None | $138.02 |
| Cloudflare | $0.15 / $0.50 | None | $138.02 |
| DeepInfra | $0.15 / $0.50 | None | $138.02 |
| Io Net | $0.15 / $0.50 | None | $138.02 |
| BaseTen | $0.15 / $0.50 | None | $138.02 |
Five of the nine already charge what most of them will be charging in a fortnight. So the 2.00x spread in that last column is not a spread between good and bad sellers, it is the promotion drawn as a picture, and OpenRouter's catalogue page reports only the bottom of it. Route to Parasail rather than Z.ai today and you pay September's price two weeks early for identical weights. Venice is the one row that will not move on September 10: its discount field reads 0, so $0.09375 and $0.3125 is its own list price rather than a markdown, and it undercuts the crowd by 37.5% with no deadline attached.
A tenth seller proved the method and then removed itself. When we pulled the endpoint list at 07:03 UTC this morning, Modal was on it at $0.149985 and $0.49995 against a discount of 0.6667. Run the same division we ran on Z.ai and that implies a list card of $0.45 and $1.50, triple what everybody else lists, which would have billed $414.06 for this run once the discount lapsed. By the time we finished writing, Modal had gone from the listing. We never found a Modal page confirming that $0.45, so it was a derivation rather than a quote either way, and the useful part is what the exercise shows: the discount field tells you which sellers have a cliff behind them, and this one decided not to find out.
Two footnotes on the Qwen side of that comparison, because the open-versus-closed line is blurrier than the table row makes it look. Alibaba did release weights this week, but for Qwen3.8-Flash-Next, a sibling whose native context is 262,144 tokens rather than the million the hosted model sells, so downloading it does not get you the thing priced here. And Alibaba runs two rate cards rather than one. Mainland China meters this model at 1 and 3 yuan per million and is reported to run substantially cheaper; the $0.16 and $0.47 in every table on this page are the international tier's own dollar prices. It is tempting to call the second a conversion of the first, and we nearly did, but the arithmetic refuses: 1 yuan is around $0.14 at any recent rate, not $0.16, and the yuan card's output-to-input ratio is exactly 3.0 where the dollar card's is 2.94. These are two separately set prices, so which one you pay is a question about which endpoint you call, not about the exchange rate.
Where the two cards genuinely disagree is the cache
Agents re-send their context. On any real long-horizon workload the cached-input line is the one that decides the invoice, and it is the line where these two models are furthest apart in ratio terms even though the numbers look almost identical today.
Take the workload our August 24 post defined, since it is cache-heavy by construction and already published: 12,000 requests in a month, 60,000 input tokens each with 70% served from cache, 12,000 output tokens each. That is 216 million fresh input tokens, 504 million cached, and 144 million output. On GLM-5.3-Flash's list card it bills $119.52. On Qwen3.8-Flash it bills $110.30. The model with the more expensive output rate wins by 8.36%, because 70% of its input costs $0.016 instead of $0.03.
Two caveats sit on that, one in each direction. Qwen3.8-Flash charges $0.20 per million to write to its cache, a quarter more than it charges to send the tokens fresh, so a prefix you write and never reuse is worse than no cache at all, by exactly $0.04 per million. That is the whole of the downside. A single read gets it back four times over: writing then reading once costs $0.216 against $0.32 for sending the same tokens twice uncached, so the break-even sits at 0.28 of a read. Any prefix used even once is ahead, which is almost every prefix an agent creates. And Z.ai publishes no cache write rate at all, listing cached storage as free for a limited time, which is a nicer deal than Alibaba's right up until the limited time is over and nobody has said when that is.
The operational difference is starker than the price one. Alibaba reports implicit caching as supported, so prefixes are matched automatically. Every GLM-5.3-Flash endpoint reports it as false, which means the 5x discount exists but you have to write code to claim it. A team that does nothing gets Qwen's cache rate and not Z.ai's.
We priced this model three days ago and missed by 8.93x
GLM-5.3-Flash is not new to this blog. It is the model that spent last week on OpenRouter as an unattributed listing called Ox Alpha, and Z.ai confirmed the connection on Wednesday. So our August 24 post about it can be marked, and it earns a split verdict.
Our identification held up. That post diffed the stealth listing against all 422 models in the catalogue, found eight of nine descriptive fields byte-identical to GLM-5.3, and argued that the ninth field ruled out the popular conclusion. The modality row said Ox Alpha took image and video and GLM-5.3 took text only, and we wrote that this pointed at a multimodal sibling rather than a rebadge, while saying plainly that no lab had confirmed anything. A multimodal sibling is exactly what it was.
Where we went wrong was the price, and not by a little. Reasoning that the model would graduate onto its family's existing card, we put the month above at $1,067.04. The card it actually arrived on bills $119.52 for the same tokens, and $59.76 while the promotion lasts. Our central estimate was 8.93x too high against the list price and 17.86x too high against the price anybody is paying this week.
That mistake is worth naming precisely, because it is easy to repeat. We treated an existing family rate card as the ceiling a new sibling would land under, and assumed the sibling would land near it. Z.ai did something the previous three GLM releases gave no reason to expect: it held $1.40 and $4.40 flat across GLM-5.1, 5.2 and 5.3, and then opened a tier underneath at a ninth of the price. Nothing in the listing we spent that post reading would have told us. The forced max-effort reasoning flag we made the centrepiece is still there and still costs what we said it costs. We just attached it to a rate card that was off by an order of magnitude.
September 10 is the only date on this page
Strip the promotion out and what is left is two closely matched models with nearly the same rate card, which is not what almost any comparison published this week says, because they read a promotional rate as a price. Both are cheap. Both take video. Both hold a million tokens. Between them, the rate cards are worth 0.22% on a measured run and the promotion is worth 100%, and only one of those two facts has an expiry date on it.
Which means the choice is not really about price at all, and the three rows that decide it are the ones we flagged at the top. If your work is cache-heavy, Alibaba's automatic 10x cache discount beats Z.ai's manual 5x, and no amount of promotional input pricing offsets it. If your output volume is the problem, only Qwen lets you cap it. If a single vendor is the problem, only GLM has nine of them and open weights under MIT, and the five sellers already charging list are a preview of what the three running the promotion do in a fortnight.
Whatever you pick, put a reminder on September 10 rather than a spreadsheet cell reading $0.075, and price your own logged token counts on both cards before then. Ours are in the pricing table with the promotional rate and the date it lapses stored as two separate facts, which is the only way we have found to keep a catalogue honest about a number that is going to change.
Six pages, two API endpoints, and one number we derived ourselves
- OpenRouter: model catalogue API - Pulled August 27, 2026, 417 records, no key required. Source of both creation timestamps (1787752741 and 1787773060, decoding to August 26 at 13:59:01 and 19:37:40 UTC, 5 hours 38 minutes apart), both context windows, both 131,072-token output caps, the text-image-video modality on both, GLM-5.3-Flash's mandatory reasoning block with max as its default effort, and Qwen3.8-Flash's non-mandatory block with reasoning-budget support
- OpenRouter: GLM-5.3-Flash endpoint detail - Pulled twice on August 27, at 07:03 and again before publishing. Ten sellers on the first pull and nine on the second, Modal having left in between. Source of every figure in the seller table, including the discount field of 0.5 on Z.AI, Novita and GMICloud, the 0.6667 Modal carried while it was listed, and the zero on Venice and on the five already charging list. Also the source of supports_implicit_caching reading false on every one of them. The equivalent Qwen3.8-Flash call returns one endpoint, Alibaba, with implicit caching true and a cache write rate of $0.20
- Z.ai: model pricing - The first-party confirmation that the discount arithmetic is right. Prints GLM-5.3-Flash at $0.15 input, $0.03 cached input and $0.50 output as the list card, the halved trio beside it, and September 9, 2026 at 24:00 UTC+8 as the end of the promotion. Cached storage is marked free for a limited time with no date given. GLM-5.3 remains $1.40, $0.26 and $4.40 on the same page
- Artificial Analysis: GLM-5.3-Flash - Intelligence Index 57, 150M output tokens across the index, $138.02 to run it, $0.09 per task, 50.2 output tokens per second, 1.47s to first token. The 420.13M input figure is ours, derived by subtracting the output leg from the published total at the list rates AA says it used. Its GLM-5.3 page supplies the comparison: index 60, 170M output tokens, $1,238.50, $0.68 per task, 85.3 tok/s, 1.57s
- Metaverse Post: Alibaba prices Qwen3.8-Flash - The $0.16 and $0.47 card, 125B total parameters plus 51B N-gram embedding parameters with 6B active per token, and a native 262,144-token context extended to a million by YaRN. Its vendor-run benchmark list, which we have not independently checked and are not comparing against GLM's: DeepSWE 1.1 58.7, SWE-bench Pro 62.5, CoWorkBench 73.9, GPQA Diamond 91.7, LiveCodeBench v6 91.9, AndroidWorld 84.5, LVBench 76.6
- TokenCost: Ox Alpha bills $0 today - Our August 24 post, the source of the 12,000-request monthly workload repriced above and of the $1,067.04 estimate this post scores. Also where the discount field first earned its place in these comparisons, reading 0 rather than 0.5
- Three things we could not establish. Modal's $0.45 and $1.50 list card is derived from its own discount field rather than read off a Modal page, and if that field is stale the $414.06 does not follow. Qwen3.8-Flash has no Artificial Analysis entry, so every figure that prices it against GLM-5.3-Flash holds GLM-5.3-Flash's measured token counts constant across both models; the two labs use different tokenizers and the real counts will differ, which moves the dollars and leaves the rate-card ratios alone. And nobody has dated the end of Z.ai's free cache storage, which is the one line on either card with no number and no deadline attached to it