Z.ai: GLM 5.2
by Z.aiZ.ai: GLM 5.2 is a large language model from Z.ai. It costs $0.683 per million input tokens and $2.15 per million output tokens. Its context window is 1.0M tokens.
- Input / 1M tokens
- $0.683
- Output / 1M tokens
- $2.15
- Cached input / 1M
- $0.127
- Context window
- 1.0M
On repeated prefixes
tokens
Who serves it cheapest
23 hosts serve GLM 5.2. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| BaiduCheapestfp8 | $0.487 | $1.53 | 1.0M | - | 99.9% |
| DeepInfrafp4 | $0.487 | $1.56 | 1.0M | - | 98.5% |
| StreamLakefp8 | $0.560 | $1.76 | 1.0M | - | 99.0% |
| Ambientfp8 | $0.600 | $2.00 | 203K | - | 97.4% |
| Novitafp8 | $0.683 | $2.15 | 1.0M | - | 100.0% |
| DigitalOcean | $0.700 | $2.20 | 262K | - | 98.9% |
| CoreWeavefp4 | $0.760 | $2.42 | 1.0M | - | 99.8% |
| AtlasCloudfp8 | $0.938 | $2.95 | 1.0M | - | 99.9% |
| Alibabafp8 | $0.966 | $3.04 | 1M | - | 99.9% |
| Inceptronfp4 | $1.10 | $2.94 | 1.0M | - | 98.5% |
| Phalafp8 | $1.26 | $3.00 | 1.0M | - | 99.9% |
| SiliconFlowfp8 | $1.19 | $3.74 | 1.0M | - | 97.8% |
| Mistral | $1.40 | $4.40 | 1.0M | - | 99.9% |
| BaseTenfp8 | $1.40 | $4.40 | 1.0M | - | 100.0% |
| Together | $1.40 | $4.40 | 512K | - | 93.9% |
| Fireworks | $1.40 | $4.40 | 1.0M | - | 99.6% |
| Venicefp8 | $1.40 | $4.40 | 1M | - | 99.8% |
| GMICloudfp8 | $1.40 | $4.40 | 1.0M | - | 99.2% |
| Parasailfp4 | $1.40 | $4.40 | 262K | - | 99.5% |
| Friendli | $1.40 | $4.40 | 1.0M | - | 99.9% |
| Cloudflare | $1.40 | $4.40 | 262K | - | 100.0% |
| Z.AIfp8 | $1.40 | $4.40 | 1.0M | - | 99.9% |
| Decartfp4 | $2.10 | $6.60 | 1.0M | - | 99.8% |
The spread between Baidu and Decart is 4.3× for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Coding index
- 68.8
- Agentic index
- 39.4
About GLM 5.2
GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...
Specifications
| Model ID | z-ai/glm-5.2 |
|---|---|
| Provider | Z.ai |
| Context window | 1.0M tokens |
| Max output | 131K tokens |
| Input modalities | text |
| Output modalities | text |
| Knowledge cutoff | - |
| Open weights | Yes - zai-org/GLM-5.2 |
| Released | June 16, 2026 |
Cheaper alternatives
Models that cost less than GLM 5.2 while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Z.ai: GLM 5.2 cost?
$0.683 per million input tokens and $2.15 per million output tokens. Cached input reads cost $0.127 per million tokens.
What is the context window of Z.ai: GLM 5.2?
1.0M tokens, with up to 131K tokens of output per request.
Which provider serves Z.ai: GLM 5.2 cheapest?
Baidu at $0.487 per million input tokens - 4.3× cheaper than Decart, the most expensive of the 23 hosts serving it.
Confirm against the source: Z.ai official pricing.