Skip to content
LLMs

LLMs that support prompt caching

inclusionAI: Ling 3.0 Flash is the cheapest option at $0.021 per million input tokens. Prompt caching bills repeated prefixes at a steep discount, which for agents and long system prompts is usually the single largest saving available. These models support it.

197 models qualify

Ranked by blended cost - input weighted 75%, output 25%.

Language models with price per million tokens and context window
#ModelBlended / 1MInput / 1MOutput / 1MContextCapabilities
1
inclusionAI logo
Ling 3.0 Flash

inclusionAI

$0.032$0.021$0.063262K
ReasoningToolsOpen weights
2$0.055$0.030$0.1301M
ReasoningToolsVision
3$0.055$0.030$0.130131K
ReasoningToolsOpen weights
4$0.057$0.050$0.080131K
ToolsOpen weights
5
Inference.net logo
Schematron V2 Turbo

Inference.net

$0.060$0.030$0.150128K
Open weights
6
Inception logo
Mercury 2.5

Inception

$0.068$0.040$0.150260K
ReasoningTools
7$0.075$0.060$0.1201.3M
ReasoningToolsOpen weights
8$0.075$0.060$0.120262K
Free tierReasoningToolsOpen weights
9$0.087$0.050$0.200262K
ReasoningToolsOpen weights
10$0.090$0.060$0.180131K
Free tierReasoningToolsVisionOpen weights
11$0.090$0.060$0.180262K
Free tierReasoningTools
12
Inference.net logo
Schematron V2 Small

Inference.net

$0.095$0.050$0.230128K
Open weights
13$0.100$0.100$0.100131K
ToolsVisionOpen weights
14$0.107$0.086$0.1711.0M
ReasoningToolsOpen weights
15
IBM Granite logo
Granite 4.2 8B

IBM Granite

$0.107$0.060$0.250131K
ReasoningToolsOpen weights
16$0.110$0.080$0.200262K
Free tierReasoningToolsOpen weights
17
Poolside logo
Laguna S 2.1

Poolside

$0.113$0.090$0.1801.0M
Free tierReasoningToolsOpen weights
18$0.125$0.100$0.2001.0M
ReasoningToolsVision
19$0.125$0.100$0.2001.0M
ReasoningToolsVision
20
ByteDance logo
UI-TARS 7B

ByteDance

$0.125$0.100$0.200128K
VisionOpen weights
21$0.131$0.075$0.300131K
ReasoningToolsOpen weights
22$0.137$0.050$0.400400K
ReasoningToolsVision
23$0.143$0.090$0.300262K
Free tierReasoningToolsVisionOpen weights
24$0.150$0.150$0.150262K
ToolsVisionOpen weights
25$0.150$0.100$0.30033K
ToolsOpen weights

Frequently asked

What is the cheapest LLM for cached prompts?

inclusionAI: Ling 3.0 Flash at $0.021 per million input tokens and $0.063 per million output tokens, with a 262K token context window.

How many models qualify for cached prompts?

197 of the 340 models tracked here meet the criteria for this list.

How much do prices vary within this category?

By roughly 950×, from Ling 3.0 Flash at the bottom to Claude Opus 4 at the top.