Z.ai: GLM 4.7
by Z.aiZ.ai: GLM 4.7 is a large language model from Z.ai. It costs $0.400 per million input tokens and $1.75 per million output tokens. Its context window is 205K tokens.
- Input / 1M tokens
- $0.400
- Output / 1M tokens
- $1.75
- Cached input / 1M
- $0.080
- Context window
- 205K
On repeated prefixes
tokens
Who serves it cheapest
7 hosts serve GLM 4.7. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| DeepInfraCheapestfp4 | $0.400 | $1.75 | 203K | - | 99.6% |
| Venicefp4 | $0.400 | $1.93 | 198K | - | 98.5% |
| AtlasCloudfp8 | $0.520 | $1.85 | 203K | - | 91.7% |
| Novitafp8 | $0.540 | $1.98 | 205K | - | 91.9% |
| $0.600 | $2.20 | 200K | - | 100.0% | |
| Z.AIfp4 | $0.600 | $2.20 | 203K | - | 82.7% |
| Mancer 2fp4 | $0.700 | $2.50 | 131K | - | 98.0% |
The spread between DeepInfra and Mancer 2 is 1.6× for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Coding index
- 45.3
About GLM 4.7
GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...
Specifications
| Model ID | z-ai/glm-4.7 |
|---|---|
| Provider | Z.ai |
| Context window | 205K tokens |
| Max output | 131K tokens |
| Input modalities | text |
| Output modalities | text |
| Knowledge cutoff | - |
| Open weights | Yes - zai-org/GLM-4.7 |
| Released | December 22, 2025 |
Cheaper alternatives
Models that cost less than GLM 4.7 while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Z.ai: GLM 4.7 cost?
$0.400 per million input tokens and $1.75 per million output tokens. Cached input reads cost $0.080 per million tokens.
What is the context window of Z.ai: GLM 4.7?
205K tokens, with up to 131K tokens of output per request.
Which provider serves Z.ai: GLM 4.7 cheapest?
DeepInfra at $0.400 per million input tokens - 1.6× cheaper than Mancer 2, the most expensive of the 7 hosts serving it.
Confirm against the source: Z.ai official pricing.