Cheapest depends on the workload
“Cheapest” depends entirely on the workload. A reasoning model is billed mostly on output; an agent is billed on call volume; a long-context job pays for the window. Each list below is ranked accordingly.
Chinese AI models
DeepSeek, Qwen, Kimi, GLM, MiniMax and the rest of the Chinese open-weight lineage. These models undercut Western equivalents by an order of magnitude and are the fastest-growing segment of the market, yet most English-language directories barely cover them.
Cheapest: Ling 3.0 Flash $0.021
121 models qualify
Models with a 1M token context window
A million tokens is roughly 750,000 words - an entire codebase or a shelf of documents in a single prompt. These are the models that offer it, and what that window costs.
Cheapest: Qwen3.7 Flash $0.030
99 models qualify
Cheapest multimodal models
Models that accept more than text and images - audio, video or documents - ranked by what the text tokens cost.
Cheapest: Qwen3.7 Flash $0.030
130 models qualify
Models with prompt caching
Prompt caching bills repeated prefixes at a steep discount, which for agents and long system prompts is usually the single largest saving available. These models support it.
Cheapest: Ling 3.0 Flash $0.021
197 models qualify
Models with a batch tier
Batch tiers trade latency for a large discount on work that does not need an answer now - backfills, evaluations, bulk classification.
Cheapest: gpt-oss-20b $0.030
77 models qualify
Cheapest frontier-class models
Models scoring in the top tier on published intelligence benchmarks, ranked by price. This is the list for the question that actually matters: the least you can pay for frontier-level capability.
Cheapest: Claude Opus 5 $5.00
3 models qualify
Cheapest LLMs for coding
Coding workloads read a lot of context and call tools constantly, so tool support and a large context window matter as much as the headline price.
Cheapest: Mistral Nemo $0.019
260 models qualify
Cheapest LLMs for agents
Agents make many calls per task, so per-token price compounds fast. These models support tool calling and structured output, the two things an agent loop cannot run without.
Cheapest: Mistral Nemo $0.019
252 models qualify
Cheapest vision models
Models that take image input alongside text, ranked by what you pay per million text tokens.
Cheapest: Qwen3.7 Flash $0.030
187 models qualify
Cheapest long-context models
Large context windows are the single most expensive thing to buy in an LLM. These models offer at least 200,000 tokens without frontier pricing.
Cheapest: Ling 3.0 Flash $0.021
218 models qualify
Cheapest reasoning models
Reasoning models emit far more output tokens than they take in, so output price dominates the bill. These are ranked with that weighting in mind.
Cheapest: Ling 3.0 Flash $0.021
218 models qualify
Cheapest open-weight models
Models whose weights you can download and run yourself. The API price here is what you pay to avoid operating the hardware.
Cheapest: Mistral Nemo $0.019
157 models qualify
Free LLM APIs
Models available at no cost, usually with rate limits and no uptime guarantee. Good for prototypes, risky for production.
Cheapest: Laguna XS 2.1 $0.060
22 models qualify