Qwen: Qwen3 32B
by QwenQwen: Qwen3 32B is a large language model from Qwen. It costs $0.080 per million input tokens and $0.280 per million output tokens. Its context window is 131K tokens.
- Input / 1M tokens
- $0.080
- Output / 1M tokens
- $0.280
- Cached input / 1M
- -
- Context window
- 131K
Not supported
tokens
Who serves it cheapest
2 hosts serve Qwen3 32B. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| DeepInfraCheapestfp8 | $0.080 | $0.280 | 41K | - | 99.6% |
| SiliconFlowfp8 | $0.140 | $0.570 | 131K | - | 99.1% |
The spread between DeepInfra and SiliconFlow is 1.9× for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Intelligence index
- 7.2
- Coding index
- 15.3
- Agentic index
- 0.9
About Qwen3 32B
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
Specifications
| Model ID | qwen/qwen3-32b |
|---|---|
| Provider | Qwen |
| Context window | 131K tokens |
| Max output | 16K tokens |
| Input modalities | text |
| Output modalities | text |
| Knowledge cutoff | 2025-03-31 |
| Open weights | Yes - Qwen/Qwen3-32B |
| Released | April 28, 2025 |
Cheaper alternatives
Models that cost less than Qwen3 32B while keeping at least half its context window and every input modality it supports.
Muse Spark 1.3 Contributor
$0.100 in · $0.200 out
about the same cheaperMuse Spark 1.2 Contributor
$0.100 in · $0.200 out
about the same cheaperUI-TARS 7B
$0.100 in · $0.200 out
about the same cheaperReka Flash 3
$0.100 in · $0.200 out
about the same cheaperQwen3 Coder 30B A3B Instruct
$0.070 in · $0.280 out
1.1× cheaperFrequently asked
How much does Qwen: Qwen3 32B cost?
$0.080 per million input tokens and $0.280 per million output tokens.
What is the context window of Qwen: Qwen3 32B?
131K tokens, with up to 16K tokens of output per request.
Which provider serves Qwen: Qwen3 32B cheapest?
DeepInfra at $0.080 per million input tokens - 1.9× cheaper than SiliconFlow, the most expensive of the 2 hosts serving it.
Confirm against the source: Qwen official pricing.