Skip to content
LLMs
NVIDIA logo

NVIDIA API pricing

6 models · US

Nemotron open models, tuned for throughput on NVIDIA inference stacks.

Cheapest input / 1M
$0.050

Nemotron 3 Nano 30B A3B

Priciest input / 1M
$0.600

Nemotron 3 Ultra

Largest context
262K

Nemotron 3 Nano 30B A3B

Models tracked
6

All NVIDIA models

Cheapest first, by blended cost.

Language models with price per million tokens and context window
#ModelBlended / 1MInput / 1MOutput / 1MContextCapabilities
1FreeFreeFree256K
Free tierReasoningToolsVisionOpen weights
2$0.087$0.050$0.200262K
ReasoningToolsOpen weights
3$0.110$0.080$0.200262K
Free tierReasoningToolsOpen weights
4$0.172$0.080$0.450262K
Free tierReasoningToolsOpen weights
5$0.200$0.200$0.200131K
Free tierReasoningVisionOpen weights
6$1.05$0.600$2.40262K
Free tierReasoningToolsOpen weights

Frequently asked

How much does the NVIDIA API cost?

NVIDIA models in this directory range from $0.050 to $0.600 per million input tokens. The cheapest is Nemotron 3 Nano 30B A3B; the most expensive is Nemotron 3 Ultra.

How many models does NVIDIA offer?

6 models are tracked here, with the largest context window being 262K tokens on Nemotron 3 Nano 30B A3B.