NVIDIA: Nemotron 3 Ultra
by NVIDIANVIDIA: Nemotron 3 Ultra is a large language model from NVIDIA. It costs $0.600 per million input tokens and $2.40 per million output tokens. Its context window is 262K tokens.
- Input / 1M tokens
- $0.600
- Output / 1M tokens
- $2.40
- Cached input / 1M
- $0.120
- Context window
- 262K
On repeated prefixes
tokens
Who serves it cheapest
3 hosts serve Nemotron 3 Ultra. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| DeepInfraCheapestfp4 | $0.500 | $2.20 | 262K | - | 99.7% |
| BaseTenfp4 | $0.600 | $2.40 | 203K | - | 100.0% |
| Venicefp8 | $0.625 | $3.13 | 256K | - | 99.8% |
The spread between DeepInfra and Venice is 1.4× for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Intelligence index
- 23.4
- Coding index
- 49.3
- Agentic index
- 21.7
About Nemotron 3 Ultra
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Specifications
| Model ID | nvidia/nemotron-3-ultra-550b-a55b |
|---|---|
| Provider | NVIDIA |
| Context window | 262K tokens |
| Max output | 183K tokens |
| Input modalities | text |
| Output modalities | text |
| Knowledge cutoff | - |
| Open weights | Yes - nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 |
| Released | June 4, 2026 |
Cheaper alternatives
Models that cost less than Nemotron 3 Ultra while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does NVIDIA: Nemotron 3 Ultra cost?
$0.600 per million input tokens and $2.40 per million output tokens. Cached input reads cost $0.120 per million tokens.
What is the context window of NVIDIA: Nemotron 3 Ultra?
262K tokens, with up to 183K tokens of output per request.
Which provider serves NVIDIA: Nemotron 3 Ultra cheapest?
DeepInfra at $0.500 per million input tokens - 1.4× cheaper than Venice, the most expensive of the 3 hosts serving it.
Confirm against the source: NVIDIA official pricing.