LLMs that support prompt caching
inclusionAI: Ling 3.0 Flash is the cheapest option at $0.021 per million input tokens. Prompt caching bills repeated prefixes at a steep discount, which for agents and long system prompts is usually the single largest saving available. These models support it.
197 models qualify
Ranked by blended cost - input weighted 75%, output 25%.
| # | Model | Blended / 1M | Input / 1M | Output / 1M | Context | Capabilities |
|---|---|---|---|---|---|---|
| 1 | Ling 3.0 Flash inclusionAI | $0.032 | $0.021 | $0.063 | 262K | ReasoningToolsOpen weights |
| 2 | Qwen3.7 Flash Qwen | $0.055 | $0.030 | $0.130 | 1M | ReasoningToolsVision |
| 3 | gpt-oss-20b OpenAI | $0.055 | $0.030 | $0.130 | 131K | ReasoningToolsOpen weights |
| 4 | Llama 3.1 8B Instruct Meta Llama | $0.057 | $0.050 | $0.080 | 131K | ToolsOpen weights |
| 5 | Schematron V2 Turbo Inference.net | $0.060 | $0.030 | $0.150 | 128K | Open weights |
| 6 | Mercury 2.5 Inception | $0.068 | $0.040 | $0.150 | 260K | ReasoningTools |
| 7 | DeepSeek V4 Flash 0731 DeepSeek | $0.075 | $0.060 | $0.120 | 1.3M | ReasoningToolsOpen weights |
| 8 | Laguna XS 2.1 Poolside | $0.075 | $0.060 | $0.120 | 262K | Free tierReasoningToolsOpen weights |
| 9 | Nemotron 3 Nano 30B A3B NVIDIA | $0.087 | $0.050 | $0.200 | 262K | ReasoningToolsOpen weights |
| 10 | Ling 3.0 Flash VL inclusionAI | $0.090 | $0.060 | $0.180 | 131K | Free tierReasoningToolsVisionOpen weights |
| 11 | Ling 3.0 Flash Fin inclusionAI | $0.090 | $0.060 | $0.180 | 262K | Free tierReasoningTools |
| 12 | Schematron V2 Small Inference.net | $0.095 | $0.050 | $0.230 | 128K | Open weights |
| 13 | Ministral 3 3B 2512 Mistral AI | $0.100 | $0.100 | $0.100 | 131K | ToolsVisionOpen weights |
| 14 | DeepSeek V4 Flash 0423 DeepSeek | $0.107 | $0.086 | $0.171 | 1.0M | ReasoningToolsOpen weights |
| 15 | Granite 4.2 8B IBM Granite | $0.107 | $0.060 | $0.250 | 131K | ReasoningToolsOpen weights |
| 16 | Nemotron 3.5 Lightning NVIDIA | $0.110 | $0.080 | $0.200 | 262K | Free tierReasoningToolsOpen weights |
| 17 | Laguna S 2.1 Poolside | $0.113 | $0.090 | $0.180 | 1.0M | Free tierReasoningToolsOpen weights |
| 18 | $0.125 | $0.100 | $0.200 | 1.0M | ReasoningToolsVision | |
| 19 | $0.125 | $0.100 | $0.200 | 1.0M | ReasoningToolsVision | |
| 20 | UI-TARS 7B ByteDance | $0.125 | $0.100 | $0.200 | 128K | VisionOpen weights |
| 21 | gpt-oss-safeguard-20b OpenAI | $0.131 | $0.075 | $0.300 | 131K | ReasoningToolsOpen weights |
| 22 | GPT-5 Nano OpenAI | $0.137 | $0.050 | $0.400 | 400K | ReasoningToolsVision |
| 23 | Gemma 4 26B A4B | $0.143 | $0.090 | $0.300 | 262K | Free tierReasoningToolsVisionOpen weights |
| 24 | Ministral 3 8B 2512 Mistral AI | $0.150 | $0.150 | $0.150 | 262K | ToolsVisionOpen weights |
| 25 | Voxtral Small 24B 2507 Mistral AI | $0.150 | $0.100 | $0.300 | 33K | ToolsOpen weights |
Frequently asked
What is the cheapest LLM for cached prompts?
inclusionAI: Ling 3.0 Flash at $0.021 per million input tokens and $0.063 per million output tokens, with a 262K token context window.
How many models qualify for cached prompts?
197 of the 340 models tracked here meet the criteria for this list.
How much do prices vary within this category?
By roughly 950×, from Ling 3.0 Flash at the bottom to Claude Opus 4 at the top.