Skip to content
LLMs
NVIDIA logo

NVIDIA: Nemotron 3.5 Lightning

by NVIDIA

NVIDIA: Nemotron 3.5 Lightning is a large language model from NVIDIA. It costs $0.080 per million input tokens and $0.200 per million output tokens. Its context window is 262K tokens.

Free tier availableReasoningTool callingStructured outputPrompt cachingOpen weights
Input / 1M tokens
$0.080
Output / 1M tokens
$0.200
Cached input / 1M
$0.040

On repeated prefixes

Context window
262K

tokens

Who serves it cheapest

4 hosts serve Nemotron 3.5 Lightning. Same weights, same API - the price difference is pure margin and routing.

Providers serving NVIDIA: Nemotron 3.5 Lightning, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
DarkbloomCheapestint4$0.065$0.180262K-99.9%
CoreWeavebf16$0.070$0.200262K-100.0%
Phala$0.080$0.200262K-99.6%
DeepInfrabf16$0.080$0.200262K-99.9%

The spread between Darkbloom and DeepInfra is 1.2× for identical weights. Quantization and context limits differ, so check both columns before switching.

Benchmarks

Independent scores published alongside the catalogue.

Intelligence index
13.6
Coding index
26.8
Agentic index
6.1

About Nemotron 3.5 Lightning

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Specifications

NVIDIA: Nemotron 3.5 Lightning specifications
Model IDnvidia/nemotron-3.5-lightning
ProviderNVIDIA
Context window262K tokens
Max output131K tokens
Input modalitiestext
Output modalitiestext
Knowledge cutoff-
Open weightsYes - nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
ReleasedAugust 11, 2026

Cheaper alternatives

Models that cost less than Nemotron 3.5 Lightning while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does NVIDIA: Nemotron 3.5 Lightning cost?

$0.080 per million input tokens and $0.200 per million output tokens. Cached input reads cost $0.040 per million tokens.

What is the context window of NVIDIA: Nemotron 3.5 Lightning?

262K tokens, with up to 131K tokens of output per request.

Which provider serves NVIDIA: Nemotron 3.5 Lightning cheapest?

Darkbloom at $0.065 per million input tokens - 1.2× cheaper than DeepInfra, the most expensive of the 4 hosts serving it.

Confirm against the source: NVIDIA official pricing.