Skip to content
LLMs
DeepSeek logo

DeepSeek: DeepSeek V4.1 Flash

by DeepSeek

DeepSeek: DeepSeek V4.1 Flash is a large language model from DeepSeek. It costs $0.150 per million input tokens and $0.600 per million output tokens. Its context window is 1.0M tokens.

ReasoningTool callingStructured outputPrompt cachingOpen weightsimage input
Input / 1M tokens
$0.150
Output / 1M tokens
$0.600
Cached input / 1M
$0.0030

On repeated prefixes

Context window
1.0M

tokens

Who serves it cheapest

17 hosts serve DeepSeek V4.1 Flash. Same weights, same API - the price difference is pure margin and routing.

Providers serving DeepSeek: DeepSeek V4.1 Flash, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
RelaceCheapestfp4$0.150$0.6001.0M-97.3%
DeepSeek$0.150$0.6001.0M-100.0%
DeepInfrafp8$0.200$0.6001.0M-94.3%
Fireworks$0.220$0.6601.0M-98.6%
Morphfp8$0.255$1.021.0M-99.2%
Io Netfp8$0.285$1.14262K-97.7%
Alibaba$0.300$1.201M-99.2%
Together$0.300$1.201.0M-96.8%
SiliconFlowfp8$0.300$1.201.0M-97.9%
Modal$0.300$1.201.0M-96.1%
Wafer$0.300$1.201.0M-98.3%
BaseTenfp8$0.300$1.201.0M-99.9%
Parasailfp8$0.300$1.201.0M-97.7%
GMICloudfp8$0.300$1.201.0M-99.9%
Novitafp8$0.300$1.201.0M-100.0%
Phala$0.345$1.381.0M-99.7%
Venicefp8$0.375$1.501M-99.5%

The spread between Relace and Venice is 2.5× for identical weights. Quantization and context limits differ, so check both columns before switching.

Benchmarks

Independent scores published alongside the catalogue.

Intelligence index
39.5

About DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

Specifications

DeepSeek: DeepSeek V4.1 Flash specifications
Model IDdeepseek/deepseek-v4.1-flash
ProviderDeepSeek
Context window1.0M tokens
Max output384K tokens
Input modalitiestext, image
Output modalitiestext
Knowledge cutoff-
Open weightsYes - deepseek-ai/DeepSeek-V4.1-Flash
ReleasedSeptember 10, 2026

Cheaper alternatives

Models that cost less than DeepSeek V4.1 Flash while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does DeepSeek: DeepSeek V4.1 Flash cost?

$0.150 per million input tokens and $0.600 per million output tokens. Cached input reads cost $0.0030 per million tokens.

What is the context window of DeepSeek: DeepSeek V4.1 Flash?

1.0M tokens, with up to 384K tokens of output per request.

Which provider serves DeepSeek: DeepSeek V4.1 Flash cheapest?

Relace at $0.150 per million input tokens - 2.5× cheaper than Venice, the most expensive of the 17 hosts serving it.

Confirm against the source: DeepSeek official pricing.