Z.ai: GLM 4.5V
by Z.aiZ.ai: GLM 4.5V is a large language model from Z.ai. It costs $0.600 per million input tokens and $1.80 per million output tokens. Its context window is 66K tokens.
- Input / 1M tokens
- $0.600
- Output / 1M tokens
- $1.80
- Cached input / 1M
- $0.110
- Context window
- 66K
On repeated prefixes
tokens
Who serves it cheapest
2 hosts serve GLM 4.5V. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| NovitaCheapestfp8 | $0.600 | $1.80 | 66K | - | 94.3% |
| Z.AIfp8 | $0.600 | $1.80 | 66K | - | 95.3% |
The spread between Novita and Z.AI is about the same for identical weights. Quantization and context limits differ, so check both columns before switching.
About GLM 4.5V
GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,...
Specifications
| Model ID | z-ai/glm-4.5v |
|---|---|
| Provider | Z.ai |
| Context window | 66K tokens |
| Max output | 16K tokens |
| Input modalities | text, image |
| Output modalities | text |
| Knowledge cutoff | 2024-12-31 |
| Open weights | Yes - zai-org/GLM-4.5V |
| Released | August 11, 2025 |
Cheaper alternatives
Models that cost less than GLM 4.5V while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Z.ai: GLM 4.5V cost?
$0.600 per million input tokens and $1.80 per million output tokens. Cached input reads cost $0.110 per million tokens.
What is the context window of Z.ai: GLM 4.5V?
66K tokens, with up to 16K tokens of output per request.
Which provider serves Z.ai: GLM 4.5V cheapest?
Novita at $0.600 per million input tokens - about the same cheaper than Z.AI, the most expensive of the 2 hosts serving it.
Confirm against the source: Z.ai official pricing.