DeepSeek: DeepSeek V4.1 Flash
by DeepSeekDeepSeek: DeepSeek V4.1 Flash is a large language model from DeepSeek. It costs $0.150 per million input tokens and $0.600 per million output tokens. Its context window is 1.0M tokens.
- Input / 1M tokens
- $0.150
- Output / 1M tokens
- $0.600
- Cached input / 1M
- $0.0030
- Context window
- 1.0M
On repeated prefixes
tokens
Who serves it cheapest
17 hosts serve DeepSeek V4.1 Flash. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| RelaceCheapestfp4 | $0.150 | $0.600 | 1.0M | - | 97.3% |
| DeepSeek | $0.150 | $0.600 | 1.0M | - | 100.0% |
| DeepInfrafp8 | $0.200 | $0.600 | 1.0M | - | 94.3% |
| Fireworks | $0.220 | $0.660 | 1.0M | - | 98.6% |
| Morphfp8 | $0.255 | $1.02 | 1.0M | - | 99.2% |
| Io Netfp8 | $0.285 | $1.14 | 262K | - | 97.7% |
| Alibaba | $0.300 | $1.20 | 1M | - | 99.2% |
| Together | $0.300 | $1.20 | 1.0M | - | 96.8% |
| SiliconFlowfp8 | $0.300 | $1.20 | 1.0M | - | 97.9% |
| Modal | $0.300 | $1.20 | 1.0M | - | 96.1% |
| Wafer | $0.300 | $1.20 | 1.0M | - | 98.3% |
| BaseTenfp8 | $0.300 | $1.20 | 1.0M | - | 99.9% |
| Parasailfp8 | $0.300 | $1.20 | 1.0M | - | 97.7% |
| GMICloudfp8 | $0.300 | $1.20 | 1.0M | - | 99.9% |
| Novitafp8 | $0.300 | $1.20 | 1.0M | - | 100.0% |
| Phala | $0.345 | $1.38 | 1.0M | - | 99.7% |
| Venicefp8 | $0.375 | $1.50 | 1M | - | 99.5% |
The spread between Relace and Venice is 2.5× for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Intelligence index
- 39.5
About DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...
Specifications
| Model ID | deepseek/deepseek-v4.1-flash |
|---|---|
| Provider | DeepSeek |
| Context window | 1.0M tokens |
| Max output | 384K tokens |
| Input modalities | text, image |
| Output modalities | text |
| Knowledge cutoff | - |
| Open weights | Yes - deepseek-ai/DeepSeek-V4.1-Flash |
| Released | September 10, 2026 |
Cheaper alternatives
Models that cost less than DeepSeek V4.1 Flash while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does DeepSeek: DeepSeek V4.1 Flash cost?
$0.150 per million input tokens and $0.600 per million output tokens. Cached input reads cost $0.0030 per million tokens.
What is the context window of DeepSeek: DeepSeek V4.1 Flash?
1.0M tokens, with up to 384K tokens of output per request.
Which provider serves DeepSeek: DeepSeek V4.1 Flash cheapest?
Relace at $0.150 per million input tokens - 2.5× cheaper than Venice, the most expensive of the 17 hosts serving it.
Confirm against the source: DeepSeek official pricing.