Skip to content
LLMs
NVIDIA logo

NVIDIA: Nemotron 3 Nano 30B A3B

by NVIDIA

NVIDIA: Nemotron 3 Nano 30B A3B is a large language model from NVIDIA. It costs $0.050 per million input tokens and $0.200 per million output tokens. Its context window is 262K tokens.

ReasoningTool callingStructured outputPrompt cachingOpen weights
Input / 1M tokens
$0.050
Output / 1M tokens
$0.200
Cached input / 1M
$0.030

On repeated prefixes

Context window
262K

tokens

Who serves it cheapest

4 hosts serve Nemotron 3 Nano 30B A3B. Same weights, same API - the price difference is pure margin and routing.

Providers serving NVIDIA: Nemotron 3 Nano 30B A3B, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
CrusoeCheapestfp8$0.050$0.200262K-99.8%
Novitafp4$0.050$0.200262K-99.9%
DeepInfrafp4$0.050$0.200262K-99.9%
Nebiusfp8$0.060$0.240262K-95.0%

The spread between Crusoe and Nebius is 1.2× for identical weights. Quantization and context limits differ, so check both columns before switching.

Benchmarks

Independent scores published alongside the catalogue.

Intelligence index
8.9
Coding index
14.4
Agentic index
1.0

About Nemotron 3 Nano 30B A3B

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...

Specifications

NVIDIA: Nemotron 3 Nano 30B A3B specifications
Model IDnvidia/nemotron-3-nano-30b-a3b
ProviderNVIDIA
Context window262K tokens
Max output236K tokens
Input modalitiestext
Output modalitiestext
Knowledge cutoff-
Open weightsYes - nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
ReleasedDecember 14, 2025

Cheaper alternatives

Models that cost less than Nemotron 3 Nano 30B A3B while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does NVIDIA: Nemotron 3 Nano 30B A3B cost?

$0.050 per million input tokens and $0.200 per million output tokens. Cached input reads cost $0.030 per million tokens.

What is the context window of NVIDIA: Nemotron 3 Nano 30B A3B?

262K tokens, with up to 236K tokens of output per request.

Which provider serves NVIDIA: Nemotron 3 Nano 30B A3B cheapest?

Crusoe at $0.050 per million input tokens - 1.2× cheaper than Nebius, the most expensive of the 4 hosts serving it.

Confirm against the source: NVIDIA official pricing.