GLM-5.3-Flash
Last verified August 27, 2026 · Zhipu pricing ↗$0.150/1M input · $0.500/1M output · 1.0M context · Zhipu
Count Tokens
—
Tokens
—
Input Cost
—
Output Cost
Estimate Monthly Cost
Monthly Cost Estimator
Quick:
<$0.0001/mo
Pick a preset above or enter custom usage
Alternatives to GLM-5.3-Flash
Pricing Details
ZhipuTHE STORED PRICE IS HALF THE LIST PRICE AND EXPIRES. Z.ai's own pricing table prints the list card as $0.15 input, $0.03 cached input, $0.50 output, with a 50% launch promotion running to 2026-09-09 24:00 UTC+8 that halves all three to $0.075/$0.015/$0.25. OpenRouter's endpoint API corroborates independently: Z.AI, Novita and GMICloud all return a discount field of 0.5, and dividing their quoted rate by (1 - discount) lands exactly on Z.ai's published list. Cached storage marked 'Limited-time Free' with NO end date given, the one line on the card with neither a number nor a deadline. THIS IS OX ALPHA: the unattributed stealth listing that appeared on OpenRouter 2026-08-20 at 20:04 UTC was a preview of this model, confirmed by Z.ai on 2026-08-26. 320B total / 18B active MoE, MIT weights on Hugging Face at zai-org/GLM-5.3-Flash. Natively multimodal, text+image+video in, text out, which is what distinguished the Ox Alpha listing from text-only GLM-5.3. REASONING IS MANDATORY and default_effort is max; supported efforts are max/high/low with no off, so reasoning tokens (billed as output) are floored rather than optional. Compare Qwen3.8-Flash, listed the same afternoon, whose reasoning block is non-mandatory with a token budget. TEN SELLERS WITHIN A DAY, NINE BY THAT AFTERNOON, and only three run the promotion: Z.AI/Novita/GMICloud at $0.075/$0.25; Parasail/Cloudflare/DeepInfra/Io Net/BaseTen already at list $0.15/$0.50; Venice at $0.09375/$0.3125 with a discount field of 0, i.e. its own genuine list, 37.5% under the crowd and NOT expiring on Sep 9. Modal was listed at 07:03 UTC 2026-08-27 at $0.149985/$0.49995 against a discount of 0.6667 (deriving a $0.45/$1.50 list we could not confirm on a Modal page) and was delisted within hours. So identical MIT weights trade at a clean 2.00x spread and OpenRouter's catalogue reports only the floor. ALL report supports_implicit_caching false, so the 5x cache discount exists but must be claimed in code; Qwen3.8-Flash's is automatic and 10x. Artificial Analysis (read 2026-08-27, billed at LIST not promo): Intelligence Index 57 against GLM-5.3's 60, 150M output tokens across the index vs the flagship's 170M, $138.02 to run it vs $1,238.50, $0.09 per task vs $0.68. Output 50.2 tok/s, TTFT 1.47s - the flagship generates 1.70x FASTER at 85.3 tok/s, so 'Flash' names the rate card, not the throughput, and the index takes roughly 1.50x the wall-clock on the cheaper model. OpenRouter also carries coding index 71.5 and agentic index 58.2 (GLM-5.3: 74.8 and 59.1). Dividing AA's published total by its published rates backs out 420.13M input tokens for that run, a 2.80:1 input-to-output ratio. Vendor-run Terminal-Bench 2.1 of 84.3 against Opus 4.8's 85.0 circulates widely; that is Z.ai's own harness, not independent, and is not stored. Estimated tokens (different tokenizer).
Input / 1M tokens
$0.15
Output / 1M tokens
$0.5
Context Window
1.0M
Max Output
131K
Price History
No price change recorded since 2026-08-27, when this site started tracking this model; it launched on 2026-08-26.
Frequently Asked Questions
Common questions about GLM-5.3-Flash pricing and usage
Read More About GLM-5.3-Flash
DeepSeek cut Flash to $0.15 this morning and will bill its Pro tier at the same rate from Monday. For the 96 hours in between, deepseek-v4-pro costs 4.4x more than the model DeepSeek says has already surpassed it.
September 10, 2026 · 13 min read
At 07:00 UTC tomorrow the cheapest model we can price goes up 5.00x, and the number that replaces it is not new. It is the card Inception has charged for Mercury 2 since March, with a fifth off the input line and nothing off the output line.
September 7, 2026 · 12 min read
Eight AI API discounts expire on a published date between September 7 and January 1. Half of them are video models nobody is tracking, the loudest one is the smallest, and a ninth that everybody lists has no end date at all.
August 28, 2026 · 14 min read