Skip to content
LLMs
Meta Llama logo

Meta: Llama 3.1 8B Instruct

by Meta Llama

Meta: Llama 3.1 8B Instruct is a large language model from Meta Llama. It costs $0.050 per million input tokens and $0.080 per million output tokens. Its context window is 131K tokens.

Tool callingStructured outputPrompt cachingOpen weights
Input / 1M tokens
$0.050
Output / 1M tokens
$0.080
Cached input / 1M
$0.025

On repeated prefixes

Context window
131K

tokens

Who serves it cheapest

5 hosts serve Llama 3.1 8B Instruct. Same weights, same API - the price difference is pure margin and routing.

Providers serving Meta: Llama 3.1 8B Instruct, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
DeepInfraCheapestfp8$0.020$0.040131K-100.0%
Novitafp8$0.020$0.05016K-95.1%
Groq$0.050$0.080131K-99.9%
Cloudflarefp8$0.152$0.28732K-97.6%
CoreWeavebf16$0.220$0.220131K-100.0%

The spread between DeepInfra and CoreWeave is 8.8× for identical weights. Quantization and context limits differ, so check both columns before switching.

Benchmarks

Independent scores published alongside the catalogue.

Coding index
5.4

About Llama 3.1 8B Instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...

Specifications

Meta: Llama 3.1 8B Instruct specifications
Model IDmeta-llama/llama-3.1-8b-instruct
ProviderMeta Llama
Context window131K tokens
Max output118K tokens
Input modalitiestext
Output modalitiestext
Knowledge cutoff2023-12-31
Open weightsYes - meta-llama/Meta-Llama-3.1-8B-Instruct
ReleasedJuly 23, 2024

Cheaper alternatives

Models that cost less than Llama 3.1 8B Instruct while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does Meta: Llama 3.1 8B Instruct cost?

$0.050 per million input tokens and $0.080 per million output tokens. Cached input reads cost $0.025 per million tokens.

What is the context window of Meta: Llama 3.1 8B Instruct?

131K tokens, with up to 118K tokens of output per request.

Which provider serves Meta: Llama 3.1 8B Instruct cheapest?

DeepInfra at $0.020 per million input tokens - 8.8× cheaper than CoreWeave, the most expensive of the 5 hosts serving it.