Z.ai: GLM 5
by Z.aiZ.ai: GLM 5 is a large language model from Z.ai. It costs $0.600 per million input tokens and $1.92 per million output tokens. Its context window is 205K tokens.
- Input / 1M tokens
- $0.600
- Output / 1M tokens
- $1.92
- Cached input / 1M
- $0.120
- Context window
- 205K
On repeated prefixes
tokens
Who serves it cheapest
8 hosts serve GLM 5. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| StreamLakeCheapestfp8 | $0.600 | $1.92 | 198K | - | 99.6% |
| GMICloudfp8 | $0.600 | $1.92 | 203K | - | 99.2% |
| Baidufp8 | $0.700 | $2.24 | 203K | - | 99.8% |
| SiliconFlowfp8 | $0.950 | $2.55 | 205K | - | 97.8% |
| Amazon Bedrock | $1.00 | $3.20 | 203K | - | 97.2% |
| Venicefp8 | $1.00 | $3.20 | 198K | - | 97.3% |
| Novitafp8 | $1.00 | $3.20 | 203K | - | 100.0% |
| Z.AIfp8 | $1.00 | $3.20 | 203K | - | 99.9% |
The spread between StreamLake and Z.AI is 1.7× for identical weights. Quantization and context limits differ, so check both columns before switching.
About GLM 5
GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading...
Specifications
| Model ID | z-ai/glm-5 |
|---|---|
| Provider | Z.ai |
| Context window | 205K tokens |
| Max output | 128K tokens |
| Input modalities | text |
| Output modalities | text |
| Knowledge cutoff | - |
| Open weights | Yes - zai-org/GLM-5 |
| Released | February 11, 2026 |
Cheaper alternatives
Models that cost less than GLM 5 while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Z.ai: GLM 5 cost?
$0.600 per million input tokens and $1.92 per million output tokens. Cached input reads cost $0.120 per million tokens.
What is the context window of Z.ai: GLM 5?
205K tokens, with up to 128K tokens of output per request.
Which provider serves Z.ai: GLM 5 cheapest?
StreamLake at $0.600 per million input tokens - 1.7× cheaper than Z.AI, the most expensive of the 8 hosts serving it.
Confirm against the source: Z.ai official pricing.