Cheapest LLMs with a 200K+ context window
inclusionAI: Ling 3.0 Flash is the cheapest option at $0.021 per million input tokens. Large context windows are the single most expensive thing to buy in an LLM. These models offer at least 200,000 tokens without frontier pricing.
218 models qualify
Ranked by blended cost - input weighted 75%, output 25%.
| # | Model | Blended / 1M | Input / 1M | Output / 1M | Context | Capabilities |
|---|---|---|---|---|---|---|
| 1 | Auto Router (Beta) OpenRouter | $-1000000.0000 | $-1000000.0000 | $-1000000.0000 | 2M | ReasoningToolsVision |
| 2 | Fusion OpenRouter | $-1000000.0000 | $-1000000.0000 | $-1000000.0000 | 1M | |
| 3 | Pareto Code Router OpenRouter | $-1000000.0000 | $-1000000.0000 | $-1000000.0000 | 2M | |
| 4 | Auto Router OpenRouter | $-1000000.0000 | $-1000000.0000 | $-1000000.0000 | 2M | ReasoningToolsVision |
| 5 | Nex-N2.5-Mini (free) Nex AGI | Free | Free | Free | 262K | Free tierReasoningToolsVisionOpen weights |
| 6 | Nex-N2.5-Pro (free) Nex AGI | Free | Free | Free | 262K | Free tierReasoningToolsVisionOpen weights |
| 7 | Ling 3.0 Flash Sante (free) inclusionAI | Free | Free | Free | 262K | Free tierReasoningTools |
| 8 | Dots3-Note Preview (free) Dots Studio | Free | Free | Free | 512K | Free tierReasoningToolsVision |
| 9 | North Mini Code (free) Cohere | Free | Free | Free | 256K | Free tierReasoningToolsOpen weights |
| 10 | Free | Free | Free | 256K | Free tierReasoningToolsVisionOpen weights | |
| 11 | Lyria 3 Pro Preview | Free | Free | Free | 1.0M | Free tierVision |
| 12 | Lyria 3 Clip Preview | Free | Free | Free | 1.0M | Free tierVision |
| 13 | Free Models Router OpenRouter | Free | Free | Free | 200K | Free tierReasoningToolsVision |
| 14 | Ling 3.0 Flash inclusionAI | $0.032 | $0.021 | $0.063 | 262K | ReasoningToolsOpen weights |
| 15 | Qwen3.7 Flash Qwen | $0.055 | $0.030 | $0.130 | 1M | ReasoningToolsVision |
| 16 | Mercury 2.5 Inception | $0.068 | $0.040 | $0.150 | 260K | ReasoningTools |
| 17 | DeepSeek V4 Flash 0731 DeepSeek | $0.075 | $0.060 | $0.120 | 1.3M | ReasoningToolsOpen weights |
| 18 | Laguna XS 2.1 Poolside | $0.075 | $0.060 | $0.120 | 262K | Free tierReasoningToolsOpen weights |
| 19 | $0.084 | $0.048 | $0.193 | 262K | ToolsOpen weights | |
| 20 | Nemotron 3 Nano 30B A3B NVIDIA | $0.087 | $0.050 | $0.200 | 262K | ReasoningToolsOpen weights |
| 21 | Ling 3.0 Flash Fin inclusionAI | $0.090 | $0.060 | $0.180 | 262K | Free tierReasoningTools |
| 22 | Nova Lite 1.0 Amazon | $0.105 | $0.060 | $0.240 | 300K | ToolsVision |
| 23 | Mistral Small 3.2 24B Mistral AI | $0.106 | $0.075 | $0.200 | 256K | ToolsVisionOpen weights |
| 24 | DeepSeek V4 Flash 0423 DeepSeek | $0.107 | $0.086 | $0.171 | 1.0M | ReasoningToolsOpen weights |
| 25 | Nemotron 3.5 Lightning NVIDIA | $0.110 | $0.080 | $0.200 | 262K | Free tierReasoningToolsOpen weights |
Frequently asked
What is the cheapest LLM for long-document work?
inclusionAI: Ling 3.0 Flash at $0.021 per million input tokens and $0.063 per million output tokens, with a 262K token context window.
How many models qualify for long-document work?
218 of the 340 models tracked here meet the criteria for this list.
How much do prices vary within this category?
By roughly 8,300×, from Ling 3.0 Flash at the bottom to o1-pro at the top.