Z.ai: GLM 4.6V
by Z.aiZ.ai: GLM 4.6V is a large language model from Z.ai. It costs $0.300 per million input tokens and $0.900 per million output tokens. Its context window is 131K tokens.
- Input / 1M tokens
- $0.300
- Output / 1M tokens
- $0.900
- Cached input / 1M
- $0.055
- Context window
- 131K
On repeated prefixes
tokens
Who serves it cheapest
2 hosts serve GLM 4.6V. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| NovitaCheapestbf16 | $0.300 | $0.900 | 131K | - | 98.0% |
| Z.AIfp8 | $0.300 | $0.900 | 131K | - | 98.0% |
The spread between Novita and Z.AI is about the same for identical weights. Quantization and context limits differ, so check both columns before switching.
About GLM 4.6V
GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...
Specifications
| Model ID | z-ai/glm-4.6v |
|---|---|
| Provider | Z.ai |
| Context window | 131K tokens |
| Max output | 33K tokens |
| Input modalities | image, text, video |
| Output modalities | text |
| Knowledge cutoff | - |
| Open weights | Yes - zai-org/GLM-4.6V |
| Released | December 8, 2025 |
Cheaper alternatives
Models that cost less than GLM 4.6V while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Z.ai: GLM 4.6V cost?
$0.300 per million input tokens and $0.900 per million output tokens. Cached input reads cost $0.055 per million tokens.
What is the context window of Z.ai: GLM 4.6V?
131K tokens, with up to 33K tokens of output per request.
Which provider serves Z.ai: GLM 4.6V cheapest?
Novita at $0.300 per million input tokens - about the same cheaper than Z.AI, the most expensive of the 2 hosts serving it.
Confirm against the source: Z.ai official pricing.