Who builds the models
52 labs, and the price range across everything each one ships. A provider trains the weights; a gateway runs them for you. Confusing the two is how people end up overpaying.
OpenAI
60 models · US
Creator of the GPT series and the de facto shape of the chat-completions API that most other providers now imitate.
Qwen
51 models · CN
Alibaba's model family, one of the most widely fine-tuned open-weight lineages on the Hub.
29 models · US
Gemini models, served through both the consumer-facing AI Studio API and Vertex AI on Google Cloud.
Mistral AI
19 models · FR
European lab shipping both open-weight and commercial models, with EU-resident inference.
Anthropic
15 models · US
Builder of the Claude family, focused on long-context reasoning, tool use and coding agents.
DeepSeek
15 models · CN
Open-weight reasoning and coding models priced far below Western equivalents, with off-peak discounts.
Z.ai
14 models · CN
The GLM series, open-weight and aggressively priced for coding workloads.
Meta Llama
8 models · US
The Llama open-weight release line - downloadable weights served by dozens of competing hosts.
MiniMax
8 models · CN
Long-context text, speech and video models with an OpenAI-compatible API.
Moonshot AI
7 models · CN
The Kimi models, strong on agentic and long-document work.
Tencent
7 models · CN
The Hunyuan models, including open-weight releases.
ByteDance Seed
6 models · CN
ByteDance's Seed research line of open-weight models.
Meta
6 models · US
Publisher of the Llama open-weight models, which most inference gateways host directly.
NVIDIA
6 models · US
Nemotron open models, tuned for throughput on NVIDIA inference stacks.
OpenRouter
6 models · US
OpenRouter's own routing models, which pick a downstream model per request.
xAI
6 models · US
The Grok family, with large context windows and native search grounding.
Amazon
5 models · US
The Nova family, served through Amazon Bedrock alongside third-party models.
Cohere
5 models · CA
Enterprise-focused Command, Embed and Rerank models built around RAG.
Perplexity
5 models · US
Sonar models with live web grounding built into the completion itself.
AionLabs
4 models
inclusionAI
4 models
Sakana
4 models
Nous Research
3 models · US
Open, community-driven fine-tunes of leading open-weight bases.
Sao10K
3 models
TheDrummer
3 models
IBM Granite
2 models · US
Apache-licensed Granite models aimed at regulated, on-premise deployments.
Inception
2 models · US
Diffusion-based language models built for very high output throughput.
Inference.net
2 models
Kwaipilot
2 models
Microsoft
2 models · US
The small, permissively licensed Phi models, plus Azure-hosted frontier models.
Morph
2 models
Nex AGI
2 models
Poolside
2 models
Reka AI
2 models · US
Natively multimodal models handling video and audio input.
Relace
2 models
StepFun
2 models
Thinking Machines
2 models
Upstage
2 models · KR
The Solar models, optimised for a strong quality-to-size ratio.
Xiaomi
2 models
Magnum v4 72B
1 model
Arcee AI
1 model · US
Small, merged and distilled models built for cost-sensitive deployment.
Baidu
1 model · CN
The ERNIE family, served through Baidu's Qianfan platform.
ByteDance
1 model · CN
The Doubao and Seed model lines.
Venice
1 model
Dots Studio
1 model
MythoMax 13B
1 model
Liquid AI
1 model · US
Liquid Foundation Models, designed for on-device and edge deployment.
Mancer
1 model
Meituan
1 model
Perceptron
1 model
ReMM SLERP 13B
1 model
Writer
1 model · US
The Palmyra models, targeted at regulated enterprise writing workflows.