GLM-5.3 bills the same $1.40 and $4.40 that GLM-5.2 and GLM-5.1 billed, and finishing a task still costs 55% more than it did. The rate card is the wrong number to be watching.
Z.ai put a per-token price on GLM-5.3 on August 18, four days after the model itself turned up reachable only through a subscription. The price is the interesting part precisely because it is not new: $1.40 input, $0.26 cached input, $4.40 output, matching GLM-5.2 on every line and GLM-5.1 before that. Three generations, one card, no movement. Then Artificial Analysis published what it costs to actually run their Intelligence Index on it, and that number went from about $0.44 on GLM-5.2 to about $0.68 on GLM-5.3. Same rates, 55% more money, because the newer model burns through materially more billable tokens reaching the same finish line: 170M output tokens across the suite against GLM-5.2's 140M, and 2.4 times the median for models in its class. The second thing worth knowing is that GLM-5.3 has exactly one seller. GLM-5.2's MIT weights put it on 21 providers, the cheapest of which will sell you input at a quarter of what Z.ai charges for GLM-5.3 today. Everything below is arithmetic on those two gaps: where GLM-5.3 genuinely undercuts Kimi K3 and where it barely does, why the answer swings between 1.15x and 3.41x depending on how much of your traffic is cached, and why the phrase "top open weights model" is doing work that the license file cannot currently back up.

Photo by Liam Scotchmer on Unsplash
The short of it
- GLM-5.3 costs $1.40 / $0.26 cached / $4.40 per million tokens. That is GLM-5.2's card and GLM-5.1's card, unchanged.
- Cost per Artificial Analysis Intelligence Index task rose to about $0.68 from about $0.44. The rates did not move, the token count did.
- It ties Kimi K3 at 60 on that index while costing $0.90 per million blended against K3's $2.31, though on cached reads alone the two are within 15% of each other.
- The weights are not out. Artificial Analysis lists GLM-5.3 as proprietary today, and Z.ai has held the release back about two weeks over its vulnerability-finding ability.
- One provider sells GLM-5.3. Twenty-one sell GLM-5.2, and the cheapest of those undercuts Z.ai's own list by about 76%. Weights, not price sheets, set what a GLM model really costs.
- No batch lane exists. Z.ai's batch inference endpoint 404s for text models, and the much-quoted 50% off-peak discount is a Coding Plan credit mechanic that never touches the metered API.
Three generations on one rate card
Z.ai's published pricing table puts GLM-5.3 and GLM-5.2 on identical numbers. Our own catalogue has carried GLM-5.1 at those same three figures since before GLM-5.2 shipped. Vendors reprice on a capability jump often enough that not repricing is itself worth a paragraph.
| Model | Input / 1M | Cached input / 1M | Output / 1M |
|---|---|---|---|
| GLM-5.3 | $1.40 | $0.26 | $4.40 |
| GLM-5.2 | $1.40 | $0.26 | $4.40 |
| GLM-5.1 | $1.40 | $0.26 | $4.40 |
| GLM-5 Turbo | $1.20 | $0.24 | $4.00 |
| GLM-5 | $1.00 | $0.20 | $3.20 |
One detail from the launch sequence is worth keeping, because Z.ai has now done it twice. GLM-5.3 arrived in mid-August reachable only through the GLM Coding Plan subscription, over the OpenAI Chat Completions-compatible protocol, with no metered rate published at all. We wrote about the same thing happening to GLM-5.2 in June, when the entire pitch was a $10 to $80 subscription and there was no token price to analyse. The gap closed in four days this time rather than the longer wait GLM-5.2 had. If you are planning around a Z.ai launch, assume the subscription comes first and the card follows.
The number that moved is not printed on the price page
Artificial Analysis runs a fixed suite and records what each model spends getting through it. On GLM-5.2 that came to roughly $0.44 per Intelligence Index task. On GLM-5.3 it is roughly $0.68. Both models bill $1.40, $0.26 and $4.40, so no part of that increase can be a price change. Every cent of it is token consumption: at identical rates, a task that costs 1.55 times as much is a task that consumed 1.55 times the billable tokens. The raw counts point the same way, with 170M output tokens across the suite against GLM-5.2's 140M, and a total index bill of $1,238.50 against $843.44. VentureBeat puts it down to the model being more verbose, which fits a release whose headline gains are on agentic and reasoning work.
Whether that trade is good depends on what you are buying. Seven points of index and a 246-point jump on GDPval-AA v2, from 1,524 to about 1,770, is a large move for a point release, and it puts GLM-5.3 second on that particular board behind Claude Opus 5. If those points translate into tasks your pipeline previously failed, paying 55% more per task to get them is straightforwardly worth it. If your workload was already inside GLM-5.2's competence, you have just signed up for a bigger bill in exchange for capability you will not use. Nothing on Z.ai's pricing page tells you which situation you are in.
One number puts the verbosity in perspective. Artificial Analysis's median output-token count for comparable models is 72M. GLM-5.3 spends 170M, which is 2.4 times that median and more than Claude Opus 5 spends to score three points higher. Verbosity is not free even when the rate card says the tokens are cheap.
If the shape feels familiar it is because we hit it a week ago from the other direction. Artificial Analysis paid 132% more per task for Grok 4.6 than for Grok 4.5 on input and output rates that had not moved either. Two unrelated vendors, two frozen cards, two materially larger bills inside a fortnight. Per-token pricing is turning into a poor proxy for what a model generation costs to use, and the gap is now wide enough that reading a static price page as a static budget is just a mistake.
Z.ai's own launch chart claims large gains on the agentic and security evals it chose, and the figures below are consistent across several outlets that saw it. We could not load Z.ai's blog directly, so treat these as vendor-reported numbers relayed secondhand rather than something we verified at source. Note also that the Terminal-Bench figure here is version 3.0, which is not the v2.1 that feeds the Artificial Analysis index. The two are different task sets and do not belong in the same column.
| Vendor-reported eval | GLM-5.2 | GLM-5.3 |
|---|---|---|
| Terminal-Bench 3.0 | 4.6 | 28.3% |
| ExploitBench | 24.4% | 54.4% |
| AutomationBench | 26.2 | 48.2% |
| CyberGym | 77.2% | 84.5% |
| DeepSWE v1.1 | 46.2 | 66.9 |
| Measure | GLM-5.2 | GLM-5.3 | Change |
|---|---|---|---|
| Price per 1M (in / cached / out) | $1.40 / $0.26 / $4.40 | $1.40 / $0.26 / $4.40 | none |
| Intelligence Index | 53 | 60 | +7 |
| Cost per Index task | ~$0.44 | ~$0.68 | +55% |
| GDPval-AA v2 Elo | 1,524 | ~1,770 | +246 |
Tied with Kimi K3, and the gap moves with your cache ratio
Both models sit at 60 on the Intelligence Index. The headline everyone picked up is that GLM-5.3 does it for $0.90 per million against Kimi K3's $2.31, which is a real figure from Artificial Analysis and reproduces exactly if you weight cached input, fresh input and output at 7:2:1. We checked: 0.7 x $0.26 plus 0.2 x $1.40 plus 0.1 x $4.40 lands on $0.902, and the same weights on K3's $0.30, $3.00 and $15.00 give $2.31 to the cent.
That single number hides a wide spread, though, and the spread is the thing you actually plan against. Line by line, K3 is 3.41 times GLM-5.3 on output, 2.14 times on fresh input, and only 1.15 times on cached reads. So the advantage is enormous if you generate a lot of text and nearly disappears if your traffic is mostly cache hits.
| Token type | GLM-5.3 | Kimi K3 | K3 as a multiple |
|---|---|---|---|
| Cached input | $0.26 | $0.30 | 1.15x |
| Fresh input | $1.40 | $3.00 | 2.14x |
| Output | $4.40 | $15.00 | 3.41x |
There is an awkward wrinkle in that. GLM-5.3 is pitched at coding and security agents, and a long-running agent loop is the most cache-dominated workload there is: the same repository context and system prompt are replayed on every turn. That is precisely the shape where the 1.15x line dominates and GLM-5.3's pricing edge is at its thinnest. The model is cheapest, relative to K3, on the work it is least specifically sold for.
Per finished task the two also sit closer than the per-token headline suggests. Artificial Analysis puts a GLM-5.3 Index task at about $0.68 against Kimi K3's $0.84, roughly 19% apart rather than the 2.5-fold gulf in the blended rate, because K3 is the less talkative of the two. Speed goes the other way and is not close: GLM-5.3 turns out about 93 output tokens a second to K3's 38, with a first token at 1.92 seconds against 2.86.
Moonshot answered the same question the opposite way
Both labs shipped a big capability jump this summer and priced it in opposite directions, which is what makes the comparison useful rather than merely tidy. Kimi K2.6 billed $0.95 and $4.00. Kimi K3 came in at $3.00 and $15.00, which is 3.16 times the old input rate and 3.75 times the old output rate. Z.ai went from GLM-5.2 to GLM-5.3 and moved nothing at all.
Neither approach is obviously right, and the per-task numbers stop this from being a simple morality tale about generous pricing. Moonshot raised the sticker and kept the model terse, so K3 finishes an Index task for $0.84. Z.ai held the sticker and let the model talk more, so GLM-5.3 finishes one for $0.68. The customer-visible difference between the two strategies is far smaller than the rate cards imply, which is a decent argument for treating any "we did not raise prices" announcement as the start of the analysis rather than the end of it.
The open weights badge has not been earned yet
A lot of the coverage describes GLM-5.3 as tying Kimi K3 for the best open weights model available. Read the Artificial Analysis entry and the open weights field says no, with the license listed as not applicable. Kimi K3's says yes, under the Kimi K3 license. The tie is real on intelligence and conditional on a release that has not happened.
Z.ai has pushed the weights back by roughly a fortnight pending a safety evaluation, and the reason reported is unusual enough to repeat: The Decoder describes the model as effective enough at finding security vulnerabilities that Z.ai is tightening controls and restricting full access to selected security partners first. We could not open Z.ai's own write-up to confirm that phrasing, so treat the specifics as secondhand. The delay itself is not in dispute, and it is a real change of posture. GLM-5.2 went out under MIT with the weights following the launch within days, which is why our catalogue can describe it as a 753B-parameter MoE. For GLM-5.3 there is currently no license to plan a self-hosting budget around, and the two-week figure is guidance rather than a committed date. We checked Hugging Face on August 22 and the newest GLM repository under Z.ai's org is still GLM-5.2. Day eight of "roughly two weeks".
Here is why that matters more than it first appears, and it is the part most of the coverage has skipped. Open weights do not just buy you the option to self-host. They create a competitive market in hosting, and that market is what actually sets the price of a GLM model. GLM-5.2 is MIT, so Artificial Analysis lists 21 providers serving it, with blended rates from $0.37 to $2.52. OpenRouter will currently sell you GLM-5.2 at $0.336 input and $1.056 output, about 76% below Z.ai's own $1.40 and $4.40. GLM-5.3 has one seller, Z.ai, at full list.
| Model | Weights | Providers | Cheapest available |
|---|---|---|---|
| GLM-5.3 | withheld | 1 | $1.40 / $4.40 (list) |
| GLM-5.2 | MIT | 21 | $0.336 / $1.056 |
| Kimi K3 | Kimi K3 license | 13 | $2.60 / $13.00 |
Sit with the top two rows for a second, because they invert the upgrade. If you are buying GLM-5.2 from the cheapest host rather than from Z.ai, moving to GLM-5.3 does not hold your price steady at all. Input goes from $0.336 to $1.40 and output from $1.056 to $4.40. Both lines land on 4.1667 times more, the same multiple to four decimal places, because the third-party rate is a flat 76.0% off Z.ai's card on input and output alike. That is before the 55% verbosity penalty lands on top of it. The frozen rate card is only frozen if Z.ai was your seller in the first place.
The same effect is visible on the other side of the comparison. Kimi K3's weights are out, so thirteen providers serve it and the cheapest asks $2.60 and $13.00 against Moonshot's own $3.00 and $15.00. Every price in the previous section was list against list, which is the fair comparison today. It stops being the fair comparison the moment GLM-5.3's weights ship, and it is already the wrong comparison if you are willing to buy last month's model from a third party.
The 50% discount everyone quotes is not for you
Two things get repeated about Z.ai pricing that do not apply to the metered API, and both will wreck a forecast if you build one on them. The first is batch. There is no batch tier for GLM-5.3, or for any Z.ai text model: the pricing page lists none, the documentation index carries async endpoints only for image and video generation, and the batch inference API reference returns a 404. Where OpenAI, Anthropic and Google will all halve your bill for accepting a delay, Z.ai sells one speed at one price.
The second is the 50% off-peak discount, which is real but lives entirely inside the GLM Coding Plan. Z.ai's developer docs define peak as Monday to Friday, 14:00 to 18:00 Singapore time, and charge model usage at half the standard credit rate outside it. That is a credit mechanic against a subscription quota. It does not touch the per-token API, where $1.40 and $4.40 apply around the clock. The subscription is worth a look on its own terms, not least because peak covers only 20 hours of the week and requests for GLM-5.2 and GLM-5.1 now route to GLM-5.3 automatically. Just do not model the two as though they share a discount.
Four workloads, priced by cache share
Monthly figures below use the published rates and nothing else. The four rows are ordered by how much of the traffic is cached, from almost all of it down to none, because that is the lever that decides how much the choice is worth. Claude Opus 5 is in the last column at $5.00, $0.50 and $25.00 as a frontier reference point.
| Monthly workload | GLM-5.3 | Kimi K3 | Claude Opus 5 |
|---|---|---|---|
| Support assistant, near-total cache900M cached, 15M fresh, 6M out | $281.40 | $405.00 | $675.00 |
| Security review agent400M cached, 40M fresh, 12M out | $212.80 | $420.00 | $700.00 |
| Nightly repo refactor120M cached, 25M fresh, 30M out | $198.20 | $561.00 | $935.00 |
| Translation queue, no caching60M fresh, 55M out | $326.00 | $1,005.00 | $1,675.00 |
Read the rows top to bottom and the pattern is the whole point. The saving against Kimi K3 runs $123.60, then $207.20, then $362.80, then $679.00, and the multiple climbs 1.44x, 1.97x, 2.83x, 3.08x in lockstep. Not one price changed between those rows. Only the cache share did. The support assistant is the awkward case: it moves the most tokens of any row and saves the least of any row, and it is also the closest thing here to what a production agent actually looks like. And none of these figures carry the verbosity penalty, because they assume you already know your token counts. If you are migrating from GLM-5.2 on the same prompts, budget for the output column growing.
Push your own numbers through the calculator rather than trusting a blended average built on somebody else's traffic shape, and the current card for every model named here is on the pricing page.
Same card, bigger bill, one seller
GLM-5.3 is a genuinely strong release at a price that has not moved in three generations, and if you are choosing between it and Kimi K3 on cost alone it wins on every line of the card. Just do not read the frozen rate card as a frozen bill. The one number Z.ai does not publish, the tokens it takes to finish your work, went up by half, and that is the number your invoice is actually made of. Anyone switching from GLM-5.2 on the strength of "same price, better model" should expect to pay more, and should decide whether the seven index points are worth it before the first invoice arrives rather than after.
The weights question is the one we would watch next, and not mainly for the self-hosting. GLM-5.2 costs $0.336 from a competitive hosting market that exists only because its weights are MIT. If GLM-5.3's land under similar terms in early September as guided, the same thing happens to it and today's $1.40 stops being the price. If the security carve-out hardens into a permanent gate, then Z.ai has quietly stopped being an open weights lab at its frontier while still being described as one, and GLM-5.3 stays a single-vendor model at single-vendor prices. That is the fork worth tracking. The rate card, for once, is the least informative thing on the page.
Sources, and what we could not check
- Z.ai: model pricing - The primary source for every rate in this post. GLM-5.3 and GLM-5.2 both at $1.40 / $0.26 / $4.40, GLM-5-Turbo at $1.20 / $0.24 / $4.00 and GLM-5 at $1.00 / $0.20 / $3.20, plus the limited-time free cached input storage. Read 2026-08-22
- Artificial Analysis: GLM-5.3 against Kimi K3 - Both at 60 on the Intelligence Index; the $0.90 and $2.31 blended figures; 93 against 38 output tokens per second; 1.92s against 2.86s to first token; and the open weights field reading no for GLM-5.3 and yes for K3. The GLM-5.3 model page adds the rank, #9 of 186, and the $0.68 per task figure
- VentureBeat: GLM-5.3 hits the API - The August 18 date for the metered card, the Coding Plan and Chat Completions-only access before it, the explicit "unchanged from GLM-5.2" wording, and the $0.68 against $0.44 per-task comparison with verbosity named as the cause
- The Decoder: top of the open model rankings, release delayed - The roughly two-week weights delay and Z.ai's stated reason, the GDPval-AA v2 move from 1,524 to 1,770, and the $0.84 figure for Kimi K3 per task. Published 2026-08-19
- TokenCost: GLM-5.2 shipped with no token price at all - June's version of the same launch pattern, when the Coding Plan was the only thing you could buy
- Z.ai: GLM Coding Plan documentation - The peak window of Monday to Friday 14:00 to 18:00 Singapore time, the 50% off-peak credit rate, and the note that GLM-5.2 and GLM-5.1 requests now route to GLM-5.3
- Moonshot: Kimi K3 pricing - $3.00 cache miss, $0.30 cache hit, $15.00 output, taxes excluded, no volume tiering documented
- What we could not verify, and where we hedged. Z.ai's own launch blog is JavaScript-rendered and we could not read it directly, so the vendor benchmark table is relayed from several outlets that did rather than checked at source. The GDPval-AA v2 figure is 1,769 on Z.ai's chart and 1,770 in Artificial Analysis's independent run; we have used the latter and flagged the former. Parameter count is disputed between 743B and 753B across reports, so we have not printed one. Kimi K3's index score here is the max reasoning setting, which is what Artificial Analysis compares; our catalogue carries an earlier default-effort reading of 57 from July and the two are not interchangeable. The two-week weights timeline is Z.ai's guidance, not a committed date, and the specific "select security partners" framing comes from The Decoder rather than from a Z.ai page we could open. Coding Plan promotional tier prices are reported inconsistently across sources and we have deliberately not quoted them. Every dollar figure in the workload table is our own arithmetic on published rates, not an invoice we have seen