Skip to content
LLMs
NVIDIA logo

NVIDIA: Nemotron 3 Super

by NVIDIA

NVIDIA: Nemotron 3 Super is a large language model from NVIDIA. It costs $0.080 per million input tokens and $0.450 per million output tokens. Its context window is 262K tokens.

Free tier availableReasoningTool callingStructured outputOpen weights
Input / 1M tokens
$0.080
Output / 1M tokens
$0.450
Cached input / 1M
-

Not supported

Context window
262K

tokens

Who serves it cheapest

2 hosts serve Nemotron 3 Super. Same weights, same API - the price difference is pure margin and routing.

Providers serving NVIDIA: Nemotron 3 Super, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
DeepInfraCheapestbf16$0.085$0.400262K-100.0%
DekaLLMfp8$0.080$0.450262K-98.5%

The spread between DeepInfra and DekaLLM is 1.1× for identical weights. Quantization and context limits differ, so check both columns before switching.

Benchmarks

Independent scores published alongside the catalogue.

Intelligence index
13.6
Coding index
37.7
Agentic index
4.1

About Nemotron 3 Super

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

Specifications

NVIDIA: Nemotron 3 Super specifications
Model IDnvidia/nemotron-3-super-120b-a12b
ProviderNVIDIA
Context window262K tokens
Max output236K tokens
Input modalitiestext
Output modalitiestext
Knowledge cutoff-
Open weightsYes - nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8
ReleasedMarch 11, 2026

Cheaper alternatives

Models that cost less than Nemotron 3 Super while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does NVIDIA: Nemotron 3 Super cost?

$0.080 per million input tokens and $0.450 per million output tokens.

What is the context window of NVIDIA: Nemotron 3 Super?

262K tokens, with up to 236K tokens of output per request.

Which provider serves NVIDIA: Nemotron 3 Super cheapest?

DeepInfra at $0.085 per million input tokens - 1.1× cheaper than DekaLLM, the most expensive of the 2 hosts serving it.

Confirm against the source: NVIDIA official pricing.