Meta: Llama 3.1 8B Instruct
by Meta LlamaMeta: Llama 3.1 8B Instruct is a large language model from Meta Llama. It costs $0.050 per million input tokens and $0.080 per million output tokens. Its context window is 131K tokens.
- Input / 1M tokens
- $0.050
- Output / 1M tokens
- $0.080
- Cached input / 1M
- $0.025
- Context window
- 131K
On repeated prefixes
tokens
Who serves it cheapest
5 hosts serve Llama 3.1 8B Instruct. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| DeepInfraCheapestfp8 | $0.020 | $0.040 | 131K | - | 100.0% |
| Novitafp8 | $0.020 | $0.050 | 16K | - | 95.1% |
| Groq | $0.050 | $0.080 | 131K | - | 99.9% |
| Cloudflarefp8 | $0.152 | $0.287 | 32K | - | 97.6% |
| CoreWeavebf16 | $0.220 | $0.220 | 131K | - | 100.0% |
The spread between DeepInfra and CoreWeave is 8.8× for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Coding index
- 5.4
About Llama 3.1 8B Instruct
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
Specifications
| Model ID | meta-llama/llama-3.1-8b-instruct |
|---|---|
| Provider | Meta Llama |
| Context window | 131K tokens |
| Max output | 118K tokens |
| Input modalities | text |
| Output modalities | text |
| Knowledge cutoff | 2023-12-31 |
| Open weights | Yes - meta-llama/Meta-Llama-3.1-8B-Instruct |
| Released | July 23, 2024 |
Cheaper alternatives
Models that cost less than Llama 3.1 8B Instruct while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Meta: Llama 3.1 8B Instruct cost?
$0.050 per million input tokens and $0.080 per million output tokens. Cached input reads cost $0.025 per million tokens.
What is the context window of Meta: Llama 3.1 8B Instruct?
131K tokens, with up to 118K tokens of output per request.
Which provider serves Meta: Llama 3.1 8B Instruct cheapest?
DeepInfra at $0.020 per million input tokens - 8.8× cheaper than CoreWeave, the most expensive of the 5 hosts serving it.