Z.ai: GLM 5.1
by Z.aiZ.ai: GLM 5.1 is a large language model from Z.ai. It costs $0.966 per million input tokens and $3.04 per million output tokens. Its context window is 205K tokens.
- Input / 1M tokens
- $0.966
- Output / 1M tokens
- $3.04
- Cached input / 1M
- $0.179
- Context window
- 205K
On repeated prefixes
tokens
Who serves it cheapest
14 hosts serve GLM 5.1. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| BaiduCheapestfp8 | $0.965 | $3.03 | 203K | - | 99.7% |
| StreamLakefp8 | $0.966 | $3.04 | 200K | - | 99.0% |
| Chutesfp8 | $0.980 | $3.08 | 203K | - | 88.1% |
| DeepInfrafp4 | $1.05 | $3.50 | 203K | - | 100.0% |
| SiliconFlowfp8 | $1.19 | $3.74 | 205K | - | 97.5% |
| AtlasCloudfp8 | $1.26 | $3.96 | 203K | - | 99.5% |
| Phala | $1.21 | $4.20 | 203K | - | 87.2% |
| Alibabafp8 | $1.33 | $4.18 | 203K | - | 100.0% |
| Novitafp8 | $1.38 | $4.40 | 205K | - | 99.6% |
| Nebiusfp8 | $1.40 | $4.40 | 203K | - | 90.5% |
| GMICloudfp8 | $1.40 | $4.40 | 203K | - | 99.8% |
| Friendli | $1.40 | $4.40 | 203K | - | 100.0% |
| Z.AIfp8 | $1.40 | $4.40 | 203K | - | 99.0% |
| Venicefp8 | $1.40 | $4.40 | 200K | - | 81.4% |
The spread between Baidu and Venice is 1.5× for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Intelligence index
- 26.4
- Coding index
- 55.8
- Agentic index
- 25.2
About GLM 5.1
GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...
Specifications
| Model ID | z-ai/glm-5.1 |
|---|---|
| Provider | Z.ai |
| Context window | 205K tokens |
| Max output | 128K tokens |
| Input modalities | text |
| Output modalities | text |
| Knowledge cutoff | - |
| Open weights | Yes - zai-org/GLM-5.1 |
| Released | April 7, 2026 |
Cheaper alternatives
Models that cost less than GLM 5.1 while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Z.ai: GLM 5.1 cost?
$0.966 per million input tokens and $3.04 per million output tokens. Cached input reads cost $0.179 per million tokens.
What is the context window of Z.ai: GLM 5.1?
205K tokens, with up to 128K tokens of output per request.
Which provider serves Z.ai: GLM 5.1 cheapest?
Baidu at $0.965 per million input tokens - 1.5× cheaper than Venice, the most expensive of the 14 hosts serving it.
Confirm against the source: Z.ai official pricing.