Skip to content
LLMs
Z.ai logo

Z.ai: GLM 4.6V

by Z.ai

Z.ai: GLM 4.6V is a large language model from Z.ai. It costs $0.300 per million input tokens and $0.900 per million output tokens. Its context window is 131K tokens.

ReasoningTool callingStructured outputPrompt cachingOpen weightsimage inputvideo input
Input / 1M tokens
$0.300
Output / 1M tokens
$0.900
Cached input / 1M
$0.055

On repeated prefixes

Context window
131K

tokens

Who serves it cheapest

2 hosts serve GLM 4.6V. Same weights, same API - the price difference is pure margin and routing.

Providers serving Z.ai: GLM 4.6V, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
NovitaCheapestbf16$0.300$0.900131K-98.0%
Z.AIfp8$0.300$0.900131K-98.0%

The spread between Novita and Z.AI is about the same for identical weights. Quantization and context limits differ, so check both columns before switching.

About GLM 4.6V

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...

Specifications

Z.ai: GLM 4.6V specifications
Model IDz-ai/glm-4.6v
ProviderZ.ai
Context window131K tokens
Max output33K tokens
Input modalitiesimage, text, video
Output modalitiestext
Knowledge cutoff-
Open weightsYes - zai-org/GLM-4.6V
ReleasedDecember 8, 2025

Cheaper alternatives

Models that cost less than GLM 4.6V while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does Z.ai: GLM 4.6V cost?

$0.300 per million input tokens and $0.900 per million output tokens. Cached input reads cost $0.055 per million tokens.

What is the context window of Z.ai: GLM 4.6V?

131K tokens, with up to 33K tokens of output per request.

Which provider serves Z.ai: GLM 4.6V cheapest?

Novita at $0.300 per million input tokens - about the same cheaper than Z.AI, the most expensive of the 2 hosts serving it.

Confirm against the source: Z.ai official pricing.