NVIDIA: Nemotron 3 Nano 30B A3B
by NVIDIANVIDIA: Nemotron 3 Nano 30B A3B is a large language model from NVIDIA. It costs $0.050 per million input tokens and $0.200 per million output tokens. Its context window is 262K tokens.
- Input / 1M tokens
- $0.050
- Output / 1M tokens
- $0.200
- Cached input / 1M
- $0.030
- Context window
- 262K
On repeated prefixes
tokens
Who serves it cheapest
4 hosts serve Nemotron 3 Nano 30B A3B. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| CrusoeCheapestfp8 | $0.050 | $0.200 | 262K | - | 99.8% |
| Novitafp4 | $0.050 | $0.200 | 262K | - | 99.9% |
| DeepInfrafp4 | $0.050 | $0.200 | 262K | - | 99.9% |
| Nebiusfp8 | $0.060 | $0.240 | 262K | - | 95.0% |
The spread between Crusoe and Nebius is 1.2× for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Intelligence index
- 8.9
- Coding index
- 14.4
- Agentic index
- 1.0
About Nemotron 3 Nano 30B A3B
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
Specifications
| Model ID | nvidia/nemotron-3-nano-30b-a3b |
|---|---|
| Provider | NVIDIA |
| Context window | 262K tokens |
| Max output | 236K tokens |
| Input modalities | text |
| Output modalities | text |
| Knowledge cutoff | - |
| Open weights | Yes - nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 |
| Released | December 14, 2025 |
Cheaper alternatives
Models that cost less than Nemotron 3 Nano 30B A3B while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does NVIDIA: Nemotron 3 Nano 30B A3B cost?
$0.050 per million input tokens and $0.200 per million output tokens. Cached input reads cost $0.030 per million tokens.
What is the context window of NVIDIA: Nemotron 3 Nano 30B A3B?
262K tokens, with up to 236K tokens of output per request.
Which provider serves NVIDIA: Nemotron 3 Nano 30B A3B cheapest?
Crusoe at $0.050 per million input tokens - 1.2× cheaper than Nebius, the most expensive of the 4 hosts serving it.
Confirm against the source: NVIDIA official pricing.