Meta: Llama 3.3 70B Instruct
by Meta LlamaMeta: Llama 3.3 70B Instruct is a large language model from Meta Llama. It costs $0.100 per million input tokens and $0.320 per million output tokens. Its context window is 131K tokens.
- Input / 1M tokens
- $0.100
- Output / 1M tokens
- $0.320
- Cached input / 1M
- -
- Context window
- 131K
Not supported
tokens
Who serves it cheapest
11 hosts serve Llama 3.3 70B Instruct. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| DeepInfraCheapestfp8 | $0.100 | $0.320 | 131K | - | 99.0% |
| Novitabf16 | $0.135 | $0.400 | 12K | - | 99.4% |
| AkashMLfp8 | $0.200 | $0.520 | 131K | - | 99.7% |
| Parasailfp8 | $0.220 | $0.500 | 131K | - | 99.8% |
| Crusoebf16 | $0.250 | $0.750 | 131K | - | 99.9% |
| SambaNova | $0.450 | $0.900 | 131K | - | 99.2% |
| Groq | $0.590 | $0.790 | 131K | - | 99.9% |
| CoreWeavefp16 | $0.710 | $0.710 | 128K | - | 99.8% |
| $0.720 | $0.720 | 128K | - | - | |
| Cloudflarefp8 | $0.293 | $2.25 | 24K | - | 98.6% |
| Together | $1.04 | $1.04 | 131K | - | 98.7% |
The spread between DeepInfra and Together is 6.7× for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Coding index
- 11.9
About Llama 3.3 70B Instruct
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
Specifications
| Model ID | meta-llama/llama-3.3-70b-instruct |
|---|---|
| Provider | Meta Llama |
| Context window | 131K tokens |
| Max output | 16K tokens |
| Input modalities | text |
| Output modalities | text |
| Knowledge cutoff | 2023-12-31 |
| Open weights | Yes - meta-llama/Llama-3.3-70B-Instruct |
| Released | December 6, 2024 |
Cheaper alternatives
Models that cost less than Llama 3.3 70B Instruct while keeping at least half its context window and every input modality it supports.
Qwen3 235B A22B Instruct 2507
$0.087 in · $0.350 out
about the same cheaperGemma 4 31B
$0.090 in · $0.340 out
about the same cheaperStep 3.5 Flash
$0.100 in · $0.300 out
about the same cheaperMinistral 3 8B 2512
$0.150 in · $0.150 out
about the same cheaperQwen3 14B
$0.120 in · $0.240 out
about the same cheaperFrequently asked
How much does Meta: Llama 3.3 70B Instruct cost?
$0.100 per million input tokens and $0.320 per million output tokens.
What is the context window of Meta: Llama 3.3 70B Instruct?
131K tokens, with up to 16K tokens of output per request.
Which provider serves Meta: Llama 3.3 70B Instruct cheapest?
DeepInfra at $0.100 per million input tokens - 6.7× cheaper than Together, the most expensive of the 11 hosts serving it.