Skip to content
LLMs
Z.ai logo

Z.ai: GLM 5.1

by Z.ai

Z.ai: GLM 5.1 is a large language model from Z.ai. It costs $0.966 per million input tokens and $3.04 per million output tokens. Its context window is 205K tokens.

ReasoningTool callingStructured outputPrompt cachingOpen weights
Input / 1M tokens
$0.966
Output / 1M tokens
$3.04
Cached input / 1M
$0.179

On repeated prefixes

Context window
205K

tokens

Who serves it cheapest

14 hosts serve GLM 5.1. Same weights, same API - the price difference is pure margin and routing.

Providers serving Z.ai: GLM 5.1, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
BaiduCheapestfp8$0.965$3.03203K-99.7%
StreamLakefp8$0.966$3.04200K-99.0%
Chutesfp8$0.980$3.08203K-88.1%
DeepInfrafp4$1.05$3.50203K-100.0%
SiliconFlowfp8$1.19$3.74205K-97.5%
AtlasCloudfp8$1.26$3.96203K-99.5%
Phala$1.21$4.20203K-87.2%
Alibabafp8$1.33$4.18203K-100.0%
Novitafp8$1.38$4.40205K-99.6%
Nebiusfp8$1.40$4.40203K-90.5%
GMICloudfp8$1.40$4.40203K-99.8%
Friendli$1.40$4.40203K-100.0%
Z.AIfp8$1.40$4.40203K-99.0%
Venicefp8$1.40$4.40200K-81.4%

The spread between Baidu and Venice is 1.5× for identical weights. Quantization and context limits differ, so check both columns before switching.

Benchmarks

Independent scores published alongside the catalogue.

Intelligence index
26.4
Coding index
55.8
Agentic index
25.2

About GLM 5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...

Specifications

Z.ai: GLM 5.1 specifications
Model IDz-ai/glm-5.1
ProviderZ.ai
Context window205K tokens
Max output128K tokens
Input modalitiestext
Output modalitiestext
Knowledge cutoff-
Open weightsYes - zai-org/GLM-5.1
ReleasedApril 7, 2026

Cheaper alternatives

Models that cost less than GLM 5.1 while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does Z.ai: GLM 5.1 cost?

$0.966 per million input tokens and $3.04 per million output tokens. Cached input reads cost $0.179 per million tokens.

What is the context window of Z.ai: GLM 5.1?

205K tokens, with up to 128K tokens of output per request.

Which provider serves Z.ai: GLM 5.1 cheapest?

Baidu at $0.965 per million input tokens - 1.5× cheaper than Venice, the most expensive of the 14 hosts serving it.

Confirm against the source: Z.ai official pricing.