Z.ai: GLM 4.7 Flash
by Z.aiZ.ai: GLM 4.7 Flash is a large language model from Z.ai. It costs $0.061 per million input tokens and $0.400 per million output tokens. Its context window is 200K tokens.
- Input / 1M tokens
- $0.061
- Output / 1M tokens
- $0.400
- Cached input / 1M
- -
- Context window
- 200K
Not supported
tokens
Who serves it cheapest
3 hosts serve GLM 4.7 Flash. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| VeniceCheapestfp8 | $0.060 | $0.400 | 128K | - | 97.2% |
| Cloudflare | $0.061 | $0.400 | 131K | - | 98.6% |
| Novitabf16 | $0.070 | $0.400 | 200K | - | 73.7% |
The spread between Venice and Novita is 1.1× for identical weights. Quantization and context limits differ, so check both columns before switching.
About GLM 4.7 Flash
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...
Specifications
| Model ID | z-ai/glm-4.7-flash |
|---|---|
| Provider | Z.ai |
| Context window | 200K tokens |
| Max output | 118K tokens |
| Input modalities | text |
| Output modalities | text |
| Knowledge cutoff | - |
| Open weights | Yes - zai-org/GLM-4.7-Flash |
| Released | January 19, 2026 |
Cheaper alternatives
Models that cost less than GLM 4.7 Flash while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Z.ai: GLM 4.7 Flash cost?
$0.061 per million input tokens and $0.400 per million output tokens.
What is the context window of Z.ai: GLM 4.7 Flash?
200K tokens, with up to 118K tokens of output per request.
Which provider serves Z.ai: GLM 4.7 Flash cheapest?
Venice at $0.060 per million input tokens - 1.1× cheaper than Novita, the most expensive of the 3 hosts serving it.
Confirm against the source: Z.ai official pricing.