Z.ai: GLM 4.6
by Z.aiZ.ai: GLM 4.6 is a large language model from Z.ai. It costs $0.430 per million input tokens and $1.75 per million output tokens. Its context window is 205K tokens.
- Input / 1M tokens
- $0.430
- Output / 1M tokens
- $1.75
- Cached input / 1M
- $0.080
- Context window
- 205K
On repeated prefixes
tokens
Who serves it cheapest
5 hosts serve GLM 4.6. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| VeniceCheapestfp4 | $0.430 | $1.75 | 198K | - | 99.9% |
| DeepInfrafp4 | $0.500 | $2.00 | 203K | - | 99.9% |
| Novitabf16 | $0.550 | $2.20 | 205K | - | 92.3% |
| AtlasCloudfp8 | $0.600 | $2.20 | 203K | - | 83.3% |
| Z.AIfp4 | $0.600 | $2.20 | 203K | - | 79.9% |
The spread between Venice and Z.AI is 1.3× for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Coding index
- 45.8
About GLM 4.6
Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...
Specifications
| Model ID | z-ai/glm-4.6 |
|---|---|
| Provider | Z.ai |
| Context window | 205K tokens |
| Max output | 16K tokens |
| Input modalities | text |
| Output modalities | text |
| Knowledge cutoff | 2025-03-31 |
| Open weights | Yes - zai-org/GLM-4.6 |
| Released | September 30, 2025 |
Cheaper alternatives
Models that cost less than GLM 4.6 while keeping at least half its context window and every input modality it supports.
Mistral Large 3 2512
$0.500 in · $1.50 out
about the same cheaperQwen3 VL 30B A3B Thinking
$0.200 in · $2.40 out
about the same cheaperQwen3 235B A22B Thinking 2507
$0.230 in · $2.30 out
about the same cheaperGLM 4.7
$0.400 in · $1.75 out
about the same cheaperQwen3.6 Plus
$0.325 in · $1.95 out
about the same cheaperFrequently asked
How much does Z.ai: GLM 4.6 cost?
$0.430 per million input tokens and $1.75 per million output tokens. Cached input reads cost $0.080 per million tokens.
What is the context window of Z.ai: GLM 4.6?
205K tokens, with up to 16K tokens of output per request.
Which provider serves Z.ai: GLM 4.6 cheapest?
Venice at $0.430 per million input tokens - 1.3× cheaper than Z.AI, the most expensive of the 5 hosts serving it.
Confirm against the source: Z.ai official pricing.