Skip to content
LLMs
Qwen logo

Qwen2.5 72B Instruct

by Qwen

Qwen2.5 72B Instruct is a large language model from Qwen. It costs $0.360 per million input tokens and $0.400 per million output tokens. Its context window is 33K tokens.

Tool callingStructured outputOpen weights
Input / 1M tokens
$0.360
Output / 1M tokens
$0.400
Cached input / 1M
-

Not supported

Context window
33K

tokens

Who serves it cheapest

2 hosts serve Qwen2.5 72B Instruct. Same weights, same API - the price difference is pure margin and routing.

Providers serving Qwen2.5 72B Instruct, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
DeepInfraCheapestfp8$0.360$0.40033K-100.0%
Novitabf16$0.380$0.40032K-97.3%

The spread between DeepInfra and Novita is about the same for identical weights. Quantization and context limits differ, so check both columns before switching.

About Qwen2.5 72B Instruct

Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...

Specifications

Qwen2.5 72B Instruct specifications
Model IDqwen/qwen-2.5-72b-instruct
ProviderQwen
Context window33K tokens
Max output16K tokens
Input modalitiestext
Output modalitiestext
Knowledge cutoff2024-06-30
Open weightsYes - Qwen/Qwen2.5-72B-Instruct
ReleasedSeptember 19, 2024

Cheaper alternatives

Models that cost less than Qwen2.5 72B Instruct while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does Qwen2.5 72B Instruct cost?

$0.360 per million input tokens and $0.400 per million output tokens.

What is the context window of Qwen2.5 72B Instruct?

33K tokens, with up to 16K tokens of output per request.

Which provider serves Qwen2.5 72B Instruct cheapest?

DeepInfra at $0.360 per million input tokens - about the same cheaper than Novita, the most expensive of the 2 hosts serving it.

Confirm against the source: Qwen official pricing.