Skip to content
LLMs
Meta Llama logo

Meta: Llama 3.3 70B Instruct

by Meta Llama

Meta: Llama 3.3 70B Instruct is a large language model from Meta Llama. It costs $0.100 per million input tokens and $0.320 per million output tokens. Its context window is 131K tokens.

Tool callingStructured outputOpen weights
Input / 1M tokens
$0.100
Output / 1M tokens
$0.320
Cached input / 1M
-

Not supported

Context window
131K

tokens

Who serves it cheapest

11 hosts serve Llama 3.3 70B Instruct. Same weights, same API - the price difference is pure margin and routing.

Providers serving Meta: Llama 3.3 70B Instruct, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
DeepInfraCheapestfp8$0.100$0.320131K-99.0%
Novitabf16$0.135$0.40012K-99.4%
AkashMLfp8$0.200$0.520131K-99.7%
Parasailfp8$0.220$0.500131K-99.8%
Crusoebf16$0.250$0.750131K-99.9%
SambaNova$0.450$0.900131K-99.2%
Groq$0.590$0.790131K-99.9%
CoreWeavefp16$0.710$0.710128K-99.8%
Google$0.720$0.720128K--
Cloudflarefp8$0.293$2.2524K-98.6%
Together$1.04$1.04131K-98.7%

The spread between DeepInfra and Together is 6.7× for identical weights. Quantization and context limits differ, so check both columns before switching.

Benchmarks

Independent scores published alongside the catalogue.

Coding index
11.9

About Llama 3.3 70B Instruct

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

Specifications

Meta: Llama 3.3 70B Instruct specifications
Model IDmeta-llama/llama-3.3-70b-instruct
ProviderMeta Llama
Context window131K tokens
Max output16K tokens
Input modalitiestext
Output modalitiestext
Knowledge cutoff2023-12-31
Open weightsYes - meta-llama/Llama-3.3-70B-Instruct
ReleasedDecember 6, 2024

Cheaper alternatives

Models that cost less than Llama 3.3 70B Instruct while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does Meta: Llama 3.3 70B Instruct cost?

$0.100 per million input tokens and $0.320 per million output tokens.

What is the context window of Meta: Llama 3.3 70B Instruct?

131K tokens, with up to 16K tokens of output per request.

Which provider serves Meta: Llama 3.3 70B Instruct cheapest?

DeepInfra at $0.100 per million input tokens - 6.7× cheaper than Together, the most expensive of the 11 hosts serving it.