Skip to content
LLMs

Cheapest depends on the workload

“Cheapest” depends entirely on the workload. A reasoning model is billed mostly on output; an agent is billed on call volume; a long-context job pays for the window. Each list below is ranked accordingly.

Chinese AI models

DeepSeek, Qwen, Kimi, GLM, MiniMax and the rest of the Chinese open-weight lineage. These models undercut Western equivalents by an order of magnitude and are the fastest-growing segment of the market, yet most English-language directories barely cover them.

Cheapest: Ling 3.0 Flash $0.021

121 models qualify

Models with a 1M token context window

A million tokens is roughly 750,000 words - an entire codebase or a shelf of documents in a single prompt. These are the models that offer it, and what that window costs.

Cheapest: Qwen3.7 Flash $0.030

99 models qualify

Cheapest multimodal models

Models that accept more than text and images - audio, video or documents - ranked by what the text tokens cost.

Cheapest: Qwen3.7 Flash $0.030

130 models qualify

Models with prompt caching

Prompt caching bills repeated prefixes at a steep discount, which for agents and long system prompts is usually the single largest saving available. These models support it.

Cheapest: Ling 3.0 Flash $0.021

197 models qualify

Models with a batch tier

Batch tiers trade latency for a large discount on work that does not need an answer now - backfills, evaluations, bulk classification.

Cheapest: gpt-oss-20b $0.030

77 models qualify

Cheapest frontier-class models

Models scoring in the top tier on published intelligence benchmarks, ranked by price. This is the list for the question that actually matters: the least you can pay for frontier-level capability.

Cheapest: Claude Opus 5 $5.00

3 models qualify

Cheapest LLMs for coding

Coding workloads read a lot of context and call tools constantly, so tool support and a large context window matter as much as the headline price.

Cheapest: Mistral Nemo $0.019

260 models qualify

Cheapest LLMs for agents

Agents make many calls per task, so per-token price compounds fast. These models support tool calling and structured output, the two things an agent loop cannot run without.

Cheapest: Mistral Nemo $0.019

252 models qualify

Cheapest vision models

Models that take image input alongside text, ranked by what you pay per million text tokens.

Cheapest: Qwen3.7 Flash $0.030

187 models qualify

Cheapest long-context models

Large context windows are the single most expensive thing to buy in an LLM. These models offer at least 200,000 tokens without frontier pricing.

Cheapest: Ling 3.0 Flash $0.021

218 models qualify

Cheapest reasoning models

Reasoning models emit far more output tokens than they take in, so output price dominates the bill. These are ranked with that weighting in mind.

Cheapest: Ling 3.0 Flash $0.021

218 models qualify

Cheapest open-weight models

Models whose weights you can download and run yourself. The API price here is what you pay to avoid operating the hardware.

Cheapest: Mistral Nemo $0.019

157 models qualify

Free LLM APIs

Models available at no cost, usually with rate limits and no uptime guarantee. Good for prototypes, risky for production.

Cheapest: Laguna XS 2.1 $0.060

22 models qualify