Skip to content
LLMs
NVIDIA logo

NVIDIA: Nemotron 3 Ultra

by NVIDIA

NVIDIA: Nemotron 3 Ultra is a large language model from NVIDIA. It costs $0.600 per million input tokens and $2.40 per million output tokens. Its context window is 262K tokens.

Free tier availableReasoningTool callingStructured outputPrompt cachingOpen weights
Input / 1M tokens
$0.600
Output / 1M tokens
$2.40
Cached input / 1M
$0.120

On repeated prefixes

Context window
262K

tokens

Who serves it cheapest

3 hosts serve Nemotron 3 Ultra. Same weights, same API - the price difference is pure margin and routing.

Providers serving NVIDIA: Nemotron 3 Ultra, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
DeepInfraCheapestfp4$0.500$2.20262K-99.7%
BaseTenfp4$0.600$2.40203K-100.0%
Venicefp8$0.625$3.13256K-99.8%

The spread between DeepInfra and Venice is 1.4× for identical weights. Quantization and context limits differ, so check both columns before switching.

Benchmarks

Independent scores published alongside the catalogue.

Intelligence index
23.4
Coding index
49.3
Agentic index
21.7

About Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Specifications

NVIDIA: Nemotron 3 Ultra specifications
Model IDnvidia/nemotron-3-ultra-550b-a55b
ProviderNVIDIA
Context window262K tokens
Max output183K tokens
Input modalitiestext
Output modalitiestext
Knowledge cutoff-
Open weightsYes - nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
ReleasedJune 4, 2026

Cheaper alternatives

Models that cost less than Nemotron 3 Ultra while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does NVIDIA: Nemotron 3 Ultra cost?

$0.600 per million input tokens and $2.40 per million output tokens. Cached input reads cost $0.120 per million tokens.

What is the context window of NVIDIA: Nemotron 3 Ultra?

262K tokens, with up to 183K tokens of output per request.

Which provider serves NVIDIA: Nemotron 3 Ultra cheapest?

DeepInfra at $0.500 per million input tokens - 1.4× cheaper than Venice, the most expensive of the 3 hosts serving it.

Confirm against the source: NVIDIA official pricing.