Meta: Llama 4 Maverick
by Meta LlamaMeta: Llama 4 Maverick is a large language model from Meta Llama. It costs $0.200 per million input tokens and $0.696 per million output tokens. Its context window is 1.0M tokens.
- Input / 1M tokens
- $0.200
- Output / 1M tokens
- $0.696
- Cached input / 1M
- -
- Context window
- 1.0M
Not supported
tokens
Who serves it cheapest
5 hosts serve Llama 4 Maverick. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| DigitalOceanCheapest | $0.200 | $0.696 | 128K | - | 99.9% |
| DeepInfrafp8 | $0.200 | $0.800 | 1.0M | - | 99.8% |
| Novitafp8 | $0.270 | $0.850 | 1.0M | - | 99.5% |
| Parasailfp8 | $0.350 | $1.00 | 524K | - | 99.5% |
| $0.350 | $1.15 | 524K | - | - |
The spread between DigitalOcean and Google is 1.7× for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Intelligence index
- 9.3
- Coding index
- 16.3
- Agentic index
- 0.6
About Llama 4 Maverick
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
Specifications
| Model ID | meta-llama/llama-4-maverick |
|---|---|
| Provider | Meta Llama |
| Context window | 1.0M tokens |
| Max output | 115K tokens |
| Input modalities | text, image |
| Output modalities | text |
| Knowledge cutoff | 2024-08-31 |
| Open weights | Yes - meta-llama/Llama-4-Maverick-17B-128E-Instruct |
| Released | April 5, 2025 |
Cheaper alternatives
Models that cost less than Llama 4 Maverick while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Meta: Llama 4 Maverick cost?
$0.200 per million input tokens and $0.696 per million output tokens.
What is the context window of Meta: Llama 4 Maverick?
1.0M tokens, with up to 115K tokens of output per request.
Which provider serves Meta: Llama 4 Maverick cheapest?
DigitalOcean at $0.200 per million input tokens - 1.7× cheaper than Google, the most expensive of the 5 hosts serving it.