GLM-5.3 Prime
Last verified October 3, 2026 · Zhipu pricing ↗$2.80/1M input · $8.80/1M output · 1.0M context · Zhipu
Count Tokens
—
Tokens
—
Input Cost
—
Output Cost
Estimate Monthly Cost
Monthly Cost Estimator
Quick:
<$0.0001/mo
Pick a preset above or enter custom usage
Alternatives to GLM-5.3 Prime
Pricing Details
ZhipuNOT a Z.ai product: Alibaba Model Studio 'Prime mode' serving Z.ai's open GLM-5.3 weights (Z.ai's pricing page has no Prime tier). OpenRouter card $2.80 input / $0.56 cache read / $8.80 output, 2x Z.ai's GLM-5.3 list ($1.40 / $4.40) and 2.35x Alibaba's own standard GLM-5.3 endpoint on OpenRouter ($1.19 / $3.74). On Alibaba's yuan price list Prime is exactly 2x standard (Beijing CNY 16 / 4 cached / 56 vs 8 / 2 / 28). Claimed 1.5-2x throughput. Measured on OpenRouter (read 2026-10-03): p50 64 tok/s, 940 ms first token (878 requests), SLOWER than Alibaba's standard GLM-5.3 endpoint (75 tok/s) and far behind third-party hosts of the same weights: Decart FP4 218 tok/s at $1.19 / $3.74, Mistral NVFP4 142 tok/s at $1.40 / $4.40, Baseten fast FP8 119 tok/s at $2.10 / $6.60. Only faster than Z.ai first-party (49 tok/s). Text only, 1M context, 131,072 max output.
Input / 1M tokens
$2.8
Output / 1M tokens
$8.8
Context Window
1.0M
Max Output
131K
Price History
Launched at current price on 2026-09-23. No price changes recorded since.
Frequently Asked Questions
Common questions about GLM-5.3 Prime pricing and usage