Z.ai: GLM 5.3 Flash
by Z.aiZ.ai: GLM 5.3 Flash is a large language model from Z.ai. It costs $0.150 per million input tokens and $0.500 per million output tokens. Its context window is 1.3M tokens.
- Input / 1M tokens
- $0.150
- Output / 1M tokens
- $0.500
- Cached input / 1M
- $0.030
- Context window
- 1.3M
On repeated prefixes
tokens
Who serves it cheapest
28 hosts serve GLM 5.3 Flash. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| DeepInfraCheapestfp4 | $0.075 | $0.250 | 1.0M | - | 98.2% |
| Relace | $0.090 | $0.300 | 1.0M | - | 97.5% |
| Morphfp8 | $0.100 | $0.350 | 1.0M | - | 28.0% |
| Wafer | $0.100 | $0.350 | 1.0M | - | 97.5% |
| StreamLakefp8 | $0.112 | $0.374 | 1.0M | - | 98.4% |
| GMICloudfp8 | $0.112 | $0.375 | 1.0M | - | 94.3% |
| Rekafp8 | $0.132 | $0.440 | 262K | - | 98.0% |
| Novitafp8 | $0.132 | $0.440 | 1.0M | - | 95.1% |
| Makora | $0.140 | $0.470 | 1.0M | - | 88.5% |
| OpenInferencefp4 | $0.150 | $0.500 | 1.0M | - | 5.0% |
| Crusoefp4 | $0.150 | $0.500 | 1.0M | - | 89.8% |
| CoreWeavefp8 | $0.150 | $0.500 | 1.0M | - | 99.7% |
| Sail Researchfp8 | $0.150 | $0.500 | 1.0M | - | 96.9% |
| AtlasCloudfp8 | $0.150 | $0.500 | 1.0M | - | 99.2% |
| Fireworks | $0.150 | $0.500 | 1.0M | - | 89.1% |
| Phalafp8 | $0.150 | $0.500 | 1.0M | - | 93.8% |
| Friendli | $0.150 | $0.500 | 1.0M | - | 96.4% |
| SiliconFlowfp8 | $0.150 | $0.500 | 1.0M | - | 83.9% |
| DigitalOcean | $0.150 | $0.500 | 1.0M | - | 98.1% |
| Together | $0.150 | $0.500 | 1.0M | - | 76.7% |
| Parasailfp8 | $0.150 | $0.500 | 1.0M | - | 86.9% |
| BaseTenfp8 | $0.150 | $0.500 | 1.0M | - | 99.8% |
| Venice | $0.150 | $0.500 | 1.0M | - | 98.4% |
| Io Netfp8 | $0.150 | $0.500 | 262K | - | 91.0% |
| Cloudflare | $0.150 | $0.500 | 1.3M | - | 97.3% |
| Z.AIfp8 | $0.150 | $0.500 | 1.0M | - | 99.3% |
| NextBitfp8 | $0.177 | $0.590 | 1.0M | - | 98.4% |
| Modalfp8 | $0.450 | $1.50 | 1.0M | - | 98.5% |
The spread between DeepInfra and Modal is 6.0× for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Intelligence index
- 41.9
- Coding index
- 71.5
- Agentic index
- 51.2
About GLM 5.3 Flash
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
Specifications
| Model ID | z-ai/glm-5.3-flash |
|---|---|
| Provider | Z.ai |
| Context window | 1.3M tokens |
| Max output | 131K tokens |
| Input modalities | text, image, video |
| Output modalities | text |
| Knowledge cutoff | - |
| Open weights | Yes - zai-org/GLM-5.3-Flash |
| Released | August 26, 2026 |
Cheaper alternatives
Models that cost less than GLM 5.3 Flash while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Z.ai: GLM 5.3 Flash cost?
$0.150 per million input tokens and $0.500 per million output tokens. Cached input reads cost $0.030 per million tokens.
What is the context window of Z.ai: GLM 5.3 Flash?
1.3M tokens, with up to 131K tokens of output per request.
Which provider serves Z.ai: GLM 5.3 Flash cheapest?
DeepInfra at $0.075 per million input tokens - 6.0× cheaper than Modal, the most expensive of the 28 hosts serving it.
Confirm against the source: Z.ai official pricing.