# Snippets AI LLM Directory > An open directory of large language models and the services that serve them. Compare real cost per million tokens, context windows and capabilities across OpenAI, Anthropic, Google, DeepSeek and 100+ inference gateways including OpenRouter - including the per-request fees and host-to-host price gaps most comparisons miss. All prices are in US dollars per 1,000,000 tokens unless stated otherwise. "Blended" cost weights input tokens at 75% and output at 25%. Data is refreshed from the live catalogue every six hours; last refreshed 2026-09-14T16:10:38.938Z. ## Key facts - Models tracked: 340 - Model providers (labs): 52 - Inference gateways and hosts: 117 - Models with a free tier: 22 - Cheapest paid model: Mistral: Mistral Nemo at $0.019 per 1M input tokens - Most expensive model: OpenAI: o1-pro at $150 per 1M input tokens ## Cheapest models - [Mistral: Mistral Nemo](https://www.getsnippets.ai/models/mistralai-mistral-nemo): $0.022 per 1M blended ($0.019 input, $0.030 output), 131K context. - [inclusionAI: Ling 3.0 Flash](https://www.getsnippets.ai/models/inclusionai-ling-3-0-flash): $0.032 per 1M blended ($0.021 input, $0.063 output), 262K context. - [IBM: Granite 4.0 Micro](https://www.getsnippets.ai/models/ibm-granite-granite-4-0-h-micro): $0.041 per 1M blended ($0.017 input, $0.112 output), 131K context. - [Sao10K: Llama 3 8B Lunaris](https://www.getsnippets.ai/models/sao10k-l3-lunaris-8b): $0.042 per 1M blended ($0.040 input, $0.050 output), 8K context. - [Qwen: Qwen3.7 Flash](https://www.getsnippets.ai/models/qwen-qwen3-7-flash): $0.055 per 1M blended ($0.030 input, $0.130 output), 1M context. - [OpenAI: gpt-oss-20b](https://www.getsnippets.ai/models/openai-gpt-oss-20b): $0.055 per 1M blended ($0.030 input, $0.130 output), 131K context. - [Mistral: Mistral Small 3](https://www.getsnippets.ai/models/mistralai-mistral-small-24b-instruct-2501): $0.057 per 1M blended ($0.050 input, $0.080 output), 33K context. - [Meta: Llama 3.1 8B Instruct](https://www.getsnippets.ai/models/meta-llama-llama-3-1-8b-instruct): $0.057 per 1M blended ($0.050 input, $0.080 output), 131K context. - [Inference.net: Schematron V2 Turbo](https://www.getsnippets.ai/models/inference-net-schematron-v2-turbo): $0.060 per 1M blended ($0.030 input, $0.150 output), 128K context. - [MythoMax 13B](https://www.getsnippets.ai/models/gryphe-mythomax-l2-13b): $0.060 per 1M blended ($0.060 input, $0.060 output), 8K context. - [Amazon: Nova Micro 1.0](https://www.getsnippets.ai/models/amazon-nova-micro-v1): $0.061 per 1M blended ($0.035 input, $0.140 output), 128K context. - [Google: Gemma 3 4B](https://www.getsnippets.ai/models/google-gemma-3-4b-it): $0.063 per 1M blended ($0.050 input, $0.100 output), 131K context. - [Cohere: Command R7B (12-2024)](https://www.getsnippets.ai/models/cohere-command-r7b-12-2024): $0.066 per 1M blended ($0.037 input, $0.150 output), 128K context. - [Inception: Mercury 2.5](https://www.getsnippets.ai/models/inception-mercury-2-5): $0.068 per 1M blended ($0.040 input, $0.150 output), 260K context. - [OpenAI: gpt-oss-120b](https://www.getsnippets.ai/models/openai-gpt-oss-120b): $0.070 per 1M blended ($0.037 input, $0.170 output), 131K context. ## Most expensive models - [OpenAI: o1-pro](https://www.getsnippets.ai/models/openai-o1-pro): $263 per 1M blended ($150 input, $600 output), 200K context. - [OpenAI: GPT-5.4 Pro](https://www.getsnippets.ai/models/openai-gpt-5-4-pro): $67.50 per 1M blended ($30.00 input, $180 output), 1.1M context. - [OpenAI: GPT-5.5 Pro](https://www.getsnippets.ai/models/openai-gpt-5-5-pro): $67.50 per 1M blended ($30.00 input, $180 output), 1.1M context. - [OpenAI: GPT-5.2 Pro](https://www.getsnippets.ai/models/openai-gpt-5-2-pro): $57.75 per 1M blended ($21.00 input, $168 output), 400K context. - [OpenAI: GPT-5 Pro](https://www.getsnippets.ai/models/openai-gpt-5-pro): $41.25 per 1M blended ($15.00 input, $120 output), 400K context. - [OpenAI: GPT-4](https://www.getsnippets.ai/models/openai-gpt-4): $37.50 per 1M blended ($30.00 input, $60.00 output), 8K context. - [OpenAI: o3 Pro](https://www.getsnippets.ai/models/openai-o3-pro): $35.00 per 1M blended ($20.00 input, $80.00 output), 200K context. - [Anthropic: Claude Opus 4](https://www.getsnippets.ai/models/anthropic-claude-opus-4): $30.00 per 1M blended ($15.00 input, $75.00 output), 200K context. - [Anthropic: Claude Opus 4.1](https://www.getsnippets.ai/models/anthropic-claude-opus-4-1): $30.00 per 1M blended ($15.00 input, $75.00 output), 200K context. - [OpenAI: o1](https://www.getsnippets.ai/models/openai-o1): $26.25 per 1M blended ($15.00 input, $60.00 output), 200K context. ## Cheapest by use case - [Chinese AI models](https://www.getsnippets.ai/cheapest/chinese-models): DeepSeek, Qwen, Kimi, GLM, MiniMax and the rest of the Chinese open-weight lineage. These models undercut Western equivalents by an order of magnitude and are the fastest-growing segment of the market, yet most English-language directories barely cover them. - [Models with a 1M token context window](https://www.getsnippets.ai/cheapest/1m-context): A million tokens is roughly 750,000 words - an entire codebase or a shelf of documents in a single prompt. These are the models that offer it, and what that window costs. - [Cheapest multimodal models](https://www.getsnippets.ai/cheapest/multimodal): Models that accept more than text and images - audio, video or documents - ranked by what the text tokens cost. - [Models with prompt caching](https://www.getsnippets.ai/cheapest/prompt-caching): Prompt caching bills repeated prefixes at a steep discount, which for agents and long system prompts is usually the single largest saving available. These models support it. - [Models with a batch tier](https://www.getsnippets.ai/cheapest/batch): Batch tiers trade latency for a large discount on work that does not need an answer now - backfills, evaluations, bulk classification. - [Cheapest frontier-class models](https://www.getsnippets.ai/cheapest/frontier): Models scoring in the top tier on published intelligence benchmarks, ranked by price. This is the list for the question that actually matters: the least you can pay for frontier-level capability. - [Cheapest LLMs for coding](https://www.getsnippets.ai/cheapest/coding): Coding workloads read a lot of context and call tools constantly, so tool support and a large context window matter as much as the headline price. - [Cheapest LLMs for agents](https://www.getsnippets.ai/cheapest/agents): Agents make many calls per task, so per-token price compounds fast. These models support tool calling and structured output, the two things an agent loop cannot run without. - [Cheapest vision models](https://www.getsnippets.ai/cheapest/vision): Models that take image input alongside text, ranked by what you pay per million text tokens. - [Cheapest long-context models](https://www.getsnippets.ai/cheapest/long-context): Large context windows are the single most expensive thing to buy in an LLM. These models offer at least 200,000 tokens without frontier pricing. - [Cheapest reasoning models](https://www.getsnippets.ai/cheapest/reasoning): Reasoning models emit far more output tokens than they take in, so output price dominates the bill. These are ranked with that weighting in mind. - [Cheapest open-weight models](https://www.getsnippets.ai/cheapest/open-weights): Models whose weights you can download and run yourself. The API price here is what you pay to avoid operating the hardware. - [Free LLM APIs](https://www.getsnippets.ai/cheapest/free): Models available at no cost, usually with rate limits and no uptime guarantee. Good for prototypes, risky for production. ## Gateways and routers (OpenRouter alternatives) - [AI/ML API](https://www.getsnippets.ai/gateways/aimlapi): Single subscription-style endpoint fronting 300+ text, image and audio models. - [Eden AI](https://www.getsnippets.ai/gateways/edenai): One API across many AI vendors, spanning language, vision, speech and OCR. - [OpenRouter](https://www.getsnippets.ai/gateways/openrouter): The reference model router: one OpenAI-compatible endpoint in front of 400+ models and 100+ hosts, with automatic price and uptime failover. - [Requesty](https://www.getsnippets.ai/gateways/requesty): Routing layer positioned as a direct OpenRouter alternative, emphasising caching and spend controls. - [Vercel AI Gateway](https://www.getsnippets.ai/gateways/vercel-ai-gateway): One endpoint and one bill across providers, wired into the Vercel AI SDK with automatic failover. ## Freshness - [New model releases](https://www.getsnippets.ai/updates): every model added to the catalogue, newest first, with launch pricing. RSS at https://www.getsnippets.ai/feed.xml. - [Discontinued gateways](https://www.getsnippets.ai/gateways/discontinued): services widely recommended as OpenRouter alternatives that have shut down, pivoted or been acquired, with the evidence and check date. ## Machine-readable data - [All models JSON](https://www.getsnippets.ai/api/models) - [All providers JSON](https://www.getsnippets.ai/api/providers) - [All gateways JSON](https://www.getsnippets.ai/api/gateways) - [API documentation](https://www.getsnippets.ai/api) - [Methodology](https://www.getsnippets.ai/methodology) ## Attribution Published by Snippets AI, Ltd. Free to cite with a link to https://www.getsnippets.ai.