Skip to content
LLMs
Z.ai logo

Z.ai: GLM 4.7 Flash

by Z.ai

Z.ai: GLM 4.7 Flash is a large language model from Z.ai. It costs $0.061 per million input tokens and $0.400 per million output tokens. Its context window is 200K tokens.

ReasoningTool callingStructured outputOpen weights
Input / 1M tokens
$0.061
Output / 1M tokens
$0.400
Cached input / 1M
-

Not supported

Context window
200K

tokens

Who serves it cheapest

3 hosts serve GLM 4.7 Flash. Same weights, same API - the price difference is pure margin and routing.

Providers serving Z.ai: GLM 4.7 Flash, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
VeniceCheapestfp8$0.060$0.400128K-97.2%
Cloudflare$0.061$0.400131K-98.6%
Novitabf16$0.070$0.400200K-73.7%

The spread between Venice and Novita is 1.1× for identical weights. Quantization and context limits differ, so check both columns before switching.

About GLM 4.7 Flash

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...

Specifications

Z.ai: GLM 4.7 Flash specifications
Model IDz-ai/glm-4.7-flash
ProviderZ.ai
Context window200K tokens
Max output118K tokens
Input modalitiestext
Output modalitiestext
Knowledge cutoff-
Open weightsYes - zai-org/GLM-4.7-Flash
ReleasedJanuary 19, 2026

Cheaper alternatives

Models that cost less than GLM 4.7 Flash while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does Z.ai: GLM 4.7 Flash cost?

$0.061 per million input tokens and $0.400 per million output tokens.

What is the context window of Z.ai: GLM 4.7 Flash?

200K tokens, with up to 118K tokens of output per request.

Which provider serves Z.ai: GLM 4.7 Flash cheapest?

Venice at $0.060 per million input tokens - 1.1× cheaper than Novita, the most expensive of the 3 hosts serving it.

Confirm against the source: Z.ai official pricing.