Qwen2.5 72B Instruct
by QwenQwen2.5 72B Instruct is a large language model from Qwen. It costs $0.360 per million input tokens and $0.400 per million output tokens. Its context window is 33K tokens.
- Input / 1M tokens
- $0.360
- Output / 1M tokens
- $0.400
- Cached input / 1M
- -
- Context window
- 33K
Not supported
tokens
Who serves it cheapest
2 hosts serve Qwen2.5 72B Instruct. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| DeepInfraCheapestfp8 | $0.360 | $0.400 | 33K | - | 100.0% |
| Novitabf16 | $0.380 | $0.400 | 32K | - | 97.3% |
The spread between DeepInfra and Novita is about the same for identical weights. Quantization and context limits differ, so check both columns before switching.
About Qwen2.5 72B Instruct
Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...
Specifications
| Model ID | qwen/qwen-2.5-72b-instruct |
|---|---|
| Provider | Qwen |
| Context window | 33K tokens |
| Max output | 16K tokens |
| Input modalities | text |
| Output modalities | text |
| Knowledge cutoff | 2024-06-30 |
| Open weights | Yes - Qwen/Qwen2.5-72B-Instruct |
| Released | September 19, 2024 |
Cheaper alternatives
Models that cost less than Qwen2.5 72B Instruct while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Qwen2.5 72B Instruct cost?
$0.360 per million input tokens and $0.400 per million output tokens.
What is the context window of Qwen2.5 72B Instruct?
33K tokens, with up to 16K tokens of output per request.
Which provider serves Qwen2.5 72B Instruct cheapest?
DeepInfra at $0.360 per million input tokens - about the same cheaper than Novita, the most expensive of the 2 hosts serving it.
Confirm against the source: Qwen official pricing.