NVIDIA API pricing
6 models · US
Nemotron open models, tuned for throughput on NVIDIA inference stacks.
- Cheapest input / 1M
- $0.050
- Priciest input / 1M
- $0.600
- Largest context
- 262K
- Models tracked
- 6
Nemotron 3 Nano 30B A3B
Nemotron 3 Ultra
Nemotron 3 Nano 30B A3B
All NVIDIA models
Cheapest first, by blended cost.
| # | Model | Blended / 1M | Input / 1M | Output / 1M | Context | Capabilities |
|---|---|---|---|---|---|---|
| 1 | Free | Free | Free | 256K | Free tierReasoningToolsVisionOpen weights | |
| 2 | Nemotron 3 Nano 30B A3B NVIDIA | $0.087 | $0.050 | $0.200 | 262K | ReasoningToolsOpen weights |
| 3 | Nemotron 3.5 Lightning NVIDIA | $0.110 | $0.080 | $0.200 | 262K | Free tierReasoningToolsOpen weights |
| 4 | Nemotron 3 Super NVIDIA | $0.172 | $0.080 | $0.450 | 262K | Free tierReasoningToolsOpen weights |
| 5 | $0.200 | $0.200 | $0.200 | 131K | Free tierReasoningVisionOpen weights | |
| 6 | Nemotron 3 Ultra NVIDIA | $1.05 | $0.600 | $2.40 | 262K | Free tierReasoningToolsOpen weights |
Frequently asked
How much does the NVIDIA API cost?
NVIDIA models in this directory range from $0.050 to $0.600 per million input tokens. The cheapest is Nemotron 3 Nano 30B A3B; the most expensive is Nemotron 3 Ultra.
How many models does NVIDIA offer?
6 models are tracked here, with the largest context window being 262K tokens on Nemotron 3 Nano 30B A3B.