Cheapest LLMs that accept images
Qwen: Qwen3.7 Flash is the cheapest option at $0.030 per million input tokens. Models that take image input alongside text, ranked by what you pay per million text tokens.
187 models qualify
Ranked by blended cost - input weighted 75%, output 25%.
| # | Model | Blended / 1M | Input / 1M | Output / 1M | Context | Capabilities |
|---|---|---|---|---|---|---|
| 1 | Auto Router (Beta) OpenRouter | $-1000000.0000 | $-1000000.0000 | $-1000000.0000 | 2M | ReasoningToolsVision |
| 2 | Auto Router OpenRouter | $-1000000.0000 | $-1000000.0000 | $-1000000.0000 | 2M | ReasoningToolsVision |
| 3 | Nex-N2.5-Mini (free) Nex AGI | Free | Free | Free | 262K | Free tierReasoningToolsVisionOpen weights |
| 4 | Nex-N2.5-Pro (free) Nex AGI | Free | Free | Free | 262K | Free tierReasoningToolsVisionOpen weights |
| 5 | Dots3-Note Preview (free) Dots Studio | Free | Free | Free | 512K | Free tierReasoningToolsVision |
| 6 | Free | Free | Free | 256K | Free tierReasoningToolsVisionOpen weights | |
| 7 | Lyria 3 Pro Preview | Free | Free | Free | 1.0M | Free tierVision |
| 8 | Lyria 3 Clip Preview | Free | Free | Free | 1.0M | Free tierVision |
| 9 | Free Models Router OpenRouter | Free | Free | Free | 200K | Free tierReasoningToolsVision |
| 10 | Qwen3.7 Flash Qwen | $0.055 | $0.030 | $0.130 | 1M | ReasoningToolsVision |
| 11 | Gemma 3 4B | $0.063 | $0.050 | $0.100 | 131K | VisionOpen weights |
| 12 | Gemma 3 12B | $0.075 | $0.050 | $0.150 | 131K | ToolsVisionOpen weights |
| 13 | Ling 3.0 Flash VL inclusionAI | $0.090 | $0.060 | $0.180 | 131K | Free tierReasoningToolsVisionOpen weights |
| 14 | Reka Edge Reka AI | $0.100 | $0.100 | $0.100 | 16K | ToolsVisionOpen weights |
| 15 | Ministral 3 3B 2512 Mistral AI | $0.100 | $0.100 | $0.100 | 131K | ToolsVisionOpen weights |
| 16 | Nova Lite 1.0 Amazon | $0.105 | $0.060 | $0.240 | 300K | ToolsVision |
| 17 | Mistral Small 3.2 24B Mistral AI | $0.106 | $0.075 | $0.200 | 256K | ToolsVisionOpen weights |
| 18 | Qwen3.5-9B Qwen | $0.112 | $0.100 | $0.150 | 262K | ReasoningToolsVisionOpen weights |
| 19 | Qwen3.5-Flash Qwen | $0.114 | $0.065 | $0.260 | 1M | ReasoningToolsVision |
| 20 | $0.125 | $0.100 | $0.200 | 1.0M | ReasoningToolsVision | |
| 21 | $0.125 | $0.100 | $0.200 | 1.0M | ReasoningToolsVision | |
| 22 | UI-TARS 7B ByteDance | $0.125 | $0.100 | $0.200 | 128K | VisionOpen weights |
| 23 | Seed 1.6 Flash ByteDance Seed | $0.131 | $0.075 | $0.300 | 262K | ReasoningToolsVision |
| 24 | GPT-5 Nano OpenAI | $0.137 | $0.050 | $0.400 | 400K | ReasoningToolsVision |
| 25 | Gemma 4 26B A4B | $0.143 | $0.090 | $0.300 | 262K | Free tierReasoningToolsVisionOpen weights |
Frequently asked
What is the cheapest LLM for image understanding?
Qwen: Qwen3.7 Flash at $0.030 per million input tokens and $0.130 per million output tokens, with a 1M token context window.
How many models qualify for image understanding?
187 of the 340 models tracked here meet the criteria for this list.
How much do prices vary within this category?
By roughly 4,800×, from Qwen3.7 Flash at the bottom to o1-pro at the top.