Qwen: Qwen3.8 2.4T A95B
by QwenQwen: Qwen3.8 2.4T A95B is a large language model from Qwen. It costs $2.00 per million input tokens and $6.00 per million output tokens. Its context window is 1.0M tokens.
- Input / 1M tokens
- $2.00
- Output / 1M tokens
- $6.00
- Cached input / 1M
- $0.250
- Context window
- 1.0M
On repeated prefixes
tokens
Who serves it cheapest
7 hosts serve Qwen3.8 2.4T A95B. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| NovitaCheapest | $2.00 | $6.00 | 1M | - | 99.8% |
| Alibaba | $2.00 | $6.00 | 1M | - | 99.7% |
| SiliconFlowfp8 | $2.00 | $6.00 | 1.0M | - | 95.7% |
| Venice | $2.00 | $6.00 | 262K | - | 97.4% |
| Modal | $2.00 | $6.00 | 1M | - | 99.9% |
| DeepInfrafp4 | $2.00 | $6.00 | 262K | - | 98.7% |
| Together | $2.00 | $6.00 | 1.0M | - | 100.0% |
The spread between Novita and Together is about the same for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Intelligence index
- 40.0
- Coding index
- 71.9
- Agentic index
- 50.4
About Qwen3.8 2.4T A95B
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...
Specifications
| Model ID | qwen/qwen3.8-2.4t-a95b |
|---|---|
| Provider | Qwen |
| Context window | 1.0M tokens |
| Max output | 131K tokens |
| Input modalities | text |
| Output modalities | text |
| Knowledge cutoff | - |
| Open weights | Yes - Qwen/Qwen3.8-2.4T-A95B |
| Released | August 12, 2026 |
Cheaper alternatives
Models that cost less than Qwen3.8 2.4T A95B while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Qwen: Qwen3.8 2.4T A95B cost?
$2.00 per million input tokens and $6.00 per million output tokens. Cached input reads cost $0.250 per million tokens.
What is the context window of Qwen: Qwen3.8 2.4T A95B?
1.0M tokens, with up to 131K tokens of output per request.
Which provider serves Qwen: Qwen3.8 2.4T A95B cheapest?
Novita at $2.00 per million input tokens - about the same cheaper than Together, the most expensive of the 7 hosts serving it.
Confirm against the source: Qwen official pricing.