Skip to content
LLMs
Z.ai logo

Z.ai: GLM 5.3 Flash

by Z.ai

Z.ai: GLM 5.3 Flash is a large language model from Z.ai. It costs $0.150 per million input tokens and $0.500 per million output tokens. Its context window is 1.3M tokens.

ReasoningTool callingStructured outputPrompt cachingBatch tierOpen weightsimage inputvideo input
Input / 1M tokens
$0.150
Output / 1M tokens
$0.500
Cached input / 1M
$0.030

On repeated prefixes

Context window
1.3M

tokens

Who serves it cheapest

28 hosts serve GLM 5.3 Flash. Same weights, same API - the price difference is pure margin and routing.

Providers serving Z.ai: GLM 5.3 Flash, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
DeepInfraCheapestfp4$0.075$0.2501.0M-98.2%
Relace$0.090$0.3001.0M-97.5%
Morphfp8$0.100$0.3501.0M-28.0%
Wafer$0.100$0.3501.0M-97.5%
StreamLakefp8$0.112$0.3741.0M-98.4%
GMICloudfp8$0.112$0.3751.0M-94.3%
Rekafp8$0.132$0.440262K-98.0%
Novitafp8$0.132$0.4401.0M-95.1%
Makora$0.140$0.4701.0M-88.5%
OpenInferencefp4$0.150$0.5001.0M-5.0%
Crusoefp4$0.150$0.5001.0M-89.8%
CoreWeavefp8$0.150$0.5001.0M-99.7%
Sail Researchfp8$0.150$0.5001.0M-96.9%
AtlasCloudfp8$0.150$0.5001.0M-99.2%
Fireworks$0.150$0.5001.0M-89.1%
Phalafp8$0.150$0.5001.0M-93.8%
Friendli$0.150$0.5001.0M-96.4%
SiliconFlowfp8$0.150$0.5001.0M-83.9%
DigitalOcean$0.150$0.5001.0M-98.1%
Together$0.150$0.5001.0M-76.7%
Parasailfp8$0.150$0.5001.0M-86.9%
BaseTenfp8$0.150$0.5001.0M-99.8%
Venice$0.150$0.5001.0M-98.4%
Io Netfp8$0.150$0.500262K-91.0%
Cloudflare$0.150$0.5001.3M-97.3%
Z.AIfp8$0.150$0.5001.0M-99.3%
NextBitfp8$0.177$0.5901.0M-98.4%
Modalfp8$0.450$1.501.0M-98.5%

The spread between DeepInfra and Modal is 6.0× for identical weights. Quantization and context limits differ, so check both columns before switching.

Benchmarks

Independent scores published alongside the catalogue.

Intelligence index
41.9
Coding index
71.5
Agentic index
51.2

About GLM 5.3 Flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Specifications

Z.ai: GLM 5.3 Flash specifications
Model IDz-ai/glm-5.3-flash
ProviderZ.ai
Context window1.3M tokens
Max output131K tokens
Input modalitiestext, image, video
Output modalitiestext
Knowledge cutoff-
Open weightsYes - zai-org/GLM-5.3-Flash
ReleasedAugust 26, 2026

Cheaper alternatives

Models that cost less than GLM 5.3 Flash while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does Z.ai: GLM 5.3 Flash cost?

$0.150 per million input tokens and $0.500 per million output tokens. Cached input reads cost $0.030 per million tokens.

What is the context window of Z.ai: GLM 5.3 Flash?

1.3M tokens, with up to 131K tokens of output per request.

Which provider serves Z.ai: GLM 5.3 Flash cheapest?

DeepInfra at $0.075 per million input tokens - 6.0× cheaper than Modal, the most expensive of the 28 hosts serving it.

Confirm against the source: Z.ai official pricing.