LLMs with discounted batch processing
OpenAI: gpt-oss-20b is the cheapest option at $0.030 per million input tokens. Batch tiers trade latency for a large discount on work that does not need an answer now - backfills, evaluations, bulk classification.
77 models qualify
Ranked by blended cost - input weighted 75%, output 25%.
| # | Model | Blended / 1M | Input / 1M | Output / 1M | Context | Capabilities |
|---|---|---|---|---|---|---|
| 1 | gpt-oss-20b OpenAI | $0.055 | $0.030 | $0.130 | 131K | ReasoningToolsOpen weights |
| 2 | gpt-oss-120b OpenAI | $0.070 | $0.037 | $0.170 | 131K | ReasoningToolsOpen weights |
| 3 | DeepSeek V4 Flash 0731 DeepSeek | $0.075 | $0.060 | $0.120 | 1.3M | ReasoningToolsOpen weights |
| 4 | Qwen3.5-9B Qwen | $0.112 | $0.100 | $0.150 | 262K | ReasoningToolsVisionOpen weights |
| 5 | GPT-5 Nano OpenAI | $0.137 | $0.050 | $0.400 | 400K | ReasoningToolsVision |
| 6 | Ministral 3 8B 2512 Mistral AI | $0.150 | $0.150 | $0.150 | 262K | ToolsVisionOpen weights |
| 7 | Gemma 4 31B | $0.152 | $0.090 | $0.340 | 262K | Free tierReasoningToolsVisionOpen weights |
| 8 | Gemini 2.5 Flash Lite | $0.175 | $0.100 | $0.400 | 1.0M | ReasoningToolsVision |
| 9 | GPT-4.1 Nano OpenAI | $0.175 | $0.100 | $0.400 | 1.0M | ToolsVision |
| 10 | GLM 5.3 Flash Z.ai | $0.237 | $0.150 | $0.500 | 1.3M | ReasoningToolsVisionOpen weights |
| 11 | Mistral Small 4 Mistral AI | $0.262 | $0.150 | $0.600 | 262K | ReasoningToolsVisionOpen weights |
| 12 | GPT-4o-mini OpenAI | $0.262 | $0.150 | $0.600 | 128K | ToolsVision |
| 13 | DeepSeek V4 Flash Vision Exp DeepSeek | $0.330 | $0.220 | $0.660 | 1.0M | ReasoningToolsVisionOpen weights |
| 14 | GPT-5.6 Luna Pro OpenAI | $0.450 | $0.200 | $1.20 | 1.1M | ReasoningToolsVision |
| 15 | GPT-5.6 Luna OpenAI | $0.450 | $0.200 | $1.20 | 1.1M | ReasoningToolsVision |
| 16 | Codestral 2508 Mistral AI | $0.450 | $0.300 | $0.900 | 256K | Tools |
| 17 | GPT-5.4 Nano OpenAI | $0.463 | $0.200 | $1.25 | 400K | ReasoningToolsVision |
| 18 | MiniMax M3 MiniMax | $0.525 | $0.300 | $1.20 | 1.0M | ReasoningToolsVisionOpen weights |
| 19 | Gemini 3.1 Flash Lite | $0.563 | $0.250 | $1.50 | 1.0M | ReasoningToolsVision |
| 20 | Muse Glimmer 30B Meta | $0.637 | $0.350 | $1.50 | 131K | ReasoningToolsVisionOpen weights |
| 21 | Inkling Small Thinking Machines | $0.637 | $0.450 | $1.20 | 1.0M | Free tierReasoningToolsVisionOpen weights |
| 22 | GPT-5 Mini OpenAI | $0.688 | $0.250 | $2.00 | 400K | ReasoningToolsVision |
| 23 | GPT-4.1 Mini OpenAI | $0.700 | $0.400 | $1.60 | 1.0M | ToolsVision |
| 24 | Mistral Large 3 2512 Mistral AI | $0.750 | $0.500 | $1.50 | 262K | ToolsVision |
| 25 | GPT-3.5 Turbo OpenAI | $0.750 | $0.500 | $1.50 | 16K | Tools |
Frequently asked
What is the cheapest LLM for asynchronous batch jobs?
OpenAI: gpt-oss-20b at $0.030 per million input tokens and $0.130 per million output tokens, with a 131K token context window.
How many models qualify for asynchronous batch jobs?
77 of the 340 models tracked here meet the criteria for this list.
How much do prices vary within this category?
By roughly 1,200×, from gpt-oss-20b at the bottom to GPT-5.4 Pro at the top.