GLM-5.3-Flash
Last verified August 27, 2026 · Zhipu pricing ↗$0.075/1M input · $0.250/1M output · 1.0M context · Zhipu
Count Tokens
—
Tokens
—
Input Cost
—
Output Cost
Estimate Monthly Cost
Monthly Cost Estimator
Quick:
<$0.0001/mo
Pick a preset above or enter custom usage
Alternatives to GLM-5.3-Flash
Pricing Details
ZhipuTHE STORED PRICE IS HALF THE LIST PRICE AND EXPIRES. Z.ai's own pricing table prints the list card as $0.15 input, $0.03 cached input, $0.50 output, with a 50% launch promotion running to 2026-09-09 24:00 UTC+8 that halves all three to $0.075/$0.015/$0.25. OpenRouter's endpoint API corroborates independently: Z.AI, Novita and GMICloud all return a discount field of 0.5, and dividing their quoted rate by (1 - discount) lands exactly on Z.ai's published list. Cached storage marked 'Limited-time Free' with NO end date given, the one line on the card with neither a number nor a deadline. THIS IS OX ALPHA: the unattributed stealth listing that appeared on OpenRouter 2026-08-20 at 20:04 UTC was a preview of this model, confirmed by Z.ai on 2026-08-26. 320B total / 18B active MoE, MIT weights on Hugging Face at zai-org/GLM-5.3-Flash. Natively multimodal, text+image+video in, text out, which is what distinguished the Ox Alpha listing from text-only GLM-5.3. REASONING IS MANDATORY and default_effort is max; supported efforts are max/high/low with no off, so reasoning tokens (billed as output) are floored rather than optional. Compare Qwen3.8-Flash, listed the same afternoon, whose reasoning block is non-mandatory with a token budget. TEN SELLERS WITHIN A DAY, NINE BY THAT AFTERNOON, and only three run the promotion: Z.AI/Novita/GMICloud at $0.075/$0.25; Parasail/Cloudflare/DeepInfra/Io Net/BaseTen already at list $0.15/$0.50; Venice at $0.09375/$0.3125 with a discount field of 0, i.e. its own genuine list, 37.5% under the crowd and NOT expiring on Sep 9. Modal was listed at 07:03 UTC 2026-08-27 at $0.149985/$0.49995 against a discount of 0.6667 (deriving a $0.45/$1.50 list we could not confirm on a Modal page) and was delisted within hours. So identical MIT weights trade at a clean 2.00x spread and OpenRouter's catalogue reports only the floor. ALL report supports_implicit_caching false, so the 5x cache discount exists but must be claimed in code; Qwen3.8-Flash's is automatic and 10x. Artificial Analysis (read 2026-08-27, billed at LIST not promo): Intelligence Index 57 against GLM-5.3's 60, 150M output tokens across the index vs the flagship's 170M, $138.02 to run it vs $1,238.50, $0.09 per task vs $0.68. Output 50.2 tok/s, TTFT 1.47s - the flagship generates 1.70x FASTER at 85.3 tok/s, so 'Flash' names the rate card, not the throughput, and the index takes roughly 1.50x the wall-clock on the cheaper model. OpenRouter also carries coding index 71.5 and agentic index 58.2 (GLM-5.3: 74.8 and 59.1). Dividing AA's published total by its published rates backs out 420.13M input tokens for that run, a 2.80:1 input-to-output ratio. Vendor-run Terminal-Bench 2.1 of 84.3 against Opus 4.8's 85.0 circulates widely; that is Z.ai's own harness, not independent, and is not stored. Estimated tokens (different tokenizer).
Input / 1M tokens
$0.075
Output / 1M tokens
$0.25
Context Window
1.0M
Max Output
131K
Launch promotion
Scheduled price change
GLM-5.3-Flash bills at $0.075 input / $0.25 output per 1M tokens through September 9, 2026. From September 10, 2026 the rate becomes $0.15 input / $0.5 output per 1M tokens, a 100% increase. Every figure on this page uses the rate in effect today.
Zhipu’s published pricing ↗Price History
Launched at current price on 2026-08-26. No price changes recorded since.
Frequently Asked Questions
Common questions about GLM-5.3-Flash pricing and usage