Z.ai: GLM 5.3
by Z.aiZ.ai: GLM 5.3 is a large language model from Z.ai. It costs $1.40 per million input tokens and $4.40 per million output tokens. Its context window is 1.3M tokens.
- Input / 1M tokens
- $1.40
- Output / 1M tokens
- $4.40
- Cached input / 1M
- $0.260
- Context window
- 1.3M
On repeated prefixes
tokens
Who serves it cheapest
26 hosts serve GLM 5.3. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| MorphCheapestfp8 | $0.920 | $3.14 | 1.0M | - | 99.7% |
| Rekafp8 | $0.936 | $3.17 | 262K | - | 99.4% |
| DigitalOcean | $0.950 | $3.40 | 1.0M | - | 98.2% |
| Novitafp8 | $1.09 | $3.43 | 1.0M | - | 99.8% |
| Phala | $1.12 | $3.52 | 1.0M | - | 99.2% |
| GMICloudfp8 | $1.12 | $3.52 | 1.0M | - | 99.4% |
| Inceptronfp4 | $1.02 | $4.11 | 1.0M | - | 99.5% |
| Decartfp4 | $1.19 | $3.74 | 1.0M | - | 99.7% |
| AkashMLfp8 | $1.17 | $3.96 | 1.0M | - | 99.5% |
| DeepInfrafp4 | $1.20 | $4.00 | 1.0M | - | 99.4% |
| Sail Researchfp8 | $1.26 | $3.95 | 1.0M | - | 99.9% |
| Friendli | $1.26 | $3.96 | 1.0M | - | 100.0% |
| Wafer | $1.19 | $4.40 | 1.0M | - | 99.4% |
| Makorafp4 | $1.35 | $4.40 | 980K | - | 97.3% |
| Baidufp8 | $1.40 | $4.40 | 1.0M | - | 99.9% |
| Crusoefp4 | $1.40 | $4.40 | 1.0M | - | 97.4% |
| Venice | $1.40 | $4.40 | 1M | - | 98.9% |
| SiliconFlowfp8 | $1.40 | $4.40 | 1.0M | - | 98.5% |
| Together | $1.40 | $4.40 | 1.0M | - | 96.5% |
| Parasailfp8 | $1.40 | $4.40 | 1.0M | - | 99.7% |
| Modal | $1.40 | $4.40 | 1.0M | - | 99.1% |
| BaseTenfp4 | $1.40 | $4.40 | 1.0M | - | 99.0% |
| Fireworks | $1.40 | $4.40 | 1.0M | - | 99.8% |
| Cloudflare | $1.40 | $4.40 | 1.3M | - | 99.9% |
| AtlasCloudfp8 | $1.40 | $4.40 | 262K | - | 99.8% |
| Z.AIfp8 | $1.40 | $4.40 | 1.0M | - | 99.7% |
The spread between Morph and Z.AI is 1.5× for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Intelligence index
- 44.9
- Coding index
- 74.8
- Agentic index
- 53.4
About GLM 5.3
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...
Specifications
| Model ID | z-ai/glm-5.3 |
|---|---|
| Provider | Z.ai |
| Context window | 1.3M tokens |
| Max output | 944K tokens |
| Input modalities | text |
| Output modalities | text |
| Knowledge cutoff | - |
| Open weights | Yes - zai-org/GLM-5.3 |
| Released | August 18, 2026 |
Cheaper alternatives
Models that cost less than GLM 5.3 while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Z.ai: GLM 5.3 cost?
$1.40 per million input tokens and $4.40 per million output tokens. Cached input reads cost $0.260 per million tokens.
What is the context window of Z.ai: GLM 5.3?
1.3M tokens, with up to 944K tokens of output per request.
Which provider serves Z.ai: GLM 5.3 cheapest?
Morph at $0.920 per million input tokens - 1.5× cheaper than Z.AI, the most expensive of the 26 hosts serving it.
Confirm against the source: Z.ai official pricing.