DeepSeek: DeepSeek V4 Flash 0731
by DeepSeekDeepSeek: DeepSeek V4 Flash 0731 is a large language model from DeepSeek. It costs $0.060 per million input tokens and $0.120 per million output tokens. Its context window is 1.3M tokens.
- Input / 1M tokens
- $0.060
- Output / 1M tokens
- $0.120
- Cached input / 1M
- $0.012
- Context window
- 1.3M
On repeated prefixes
tokens
Who serves it cheapest
27 hosts serve DeepSeek V4 Flash 0731. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| OpenInferenceCheapestfp8 | $0.040 | $0.100 | 1.0M | - | 98.8% |
| Relacefp4 | $0.060 | $0.120 | 1.0M | - | 99.5% |
| StreamLakefp8 | $0.057 | $0.172 | 1.0M | - | 97.7% |
| DeepInfrafp8 | $0.060 | $0.180 | 1.0M | - | 99.7% |
| Inceptronfp4 | $0.064 | $0.173 | 1.0M | - | 99.7% |
| Makora | $0.090 | $0.195 | 1M | - | 98.0% |
| Wafer | $0.100 | $0.250 | 1.0M | - | 99.3% |
| Sail Researchfp4 | $0.074 | $0.342 | 1.0M | - | 99.5% |
| DigitalOcean | $0.119 | $0.238 | 1.0M | - | 99.8% |
| BaseTenfp8 | $0.130 | $0.260 | 1.0M | - | 100.0% |
| CoreWeavefp8 | $0.130 | $0.280 | 262K | - | 100.0% |
| Together | $0.140 | $0.280 | 1.0M | - | 97.6% |
| Parasailfp8 | $0.140 | $0.280 | 1.0M | - | 99.4% |
| Morphbf16 | $0.123 | $0.348 | 1.0M | - | 99.1% |
| Venice | $0.175 | $0.350 | 1M | - | 99.1% |
| Rekafp4 | $0.110 | $0.660 | 262K | - | 99.6% |
| Alibaba | $0.176 | $0.528 | 1M | - | 100.0% |
| Mancer 2fp8 | $0.200 | $0.600 | 1.0M | - | 98.1% |
| Fireworks | $0.220 | $0.660 | 1.0M | - | 98.2% |
| SiliconFlowfp8 | $0.220 | $0.660 | 1.0M | - | 93.5% |
| GMICloudfp8 | $0.286 | $0.858 | 1.0M | - | 99.9% |
| Phala | $0.308 | $0.924 | 1.0M | - | 99.8% |
| NextBitfp8 | $0.352 | $1.06 | 1.0M | - | 99.7% |
| Novitafp8 | $0.409 | $1.23 | 1.0M | - | 100.0% |
| Baidufp8 | $0.440 | $1.32 | 1.0M | - | 99.8% |
| AtlasCloudfp4 | $0.440 | $1.32 | 1.0M | - | 99.9% |
| Cloudflare | $0.440 | $1.32 | 1.3M | - | 100.0% |
The spread between OpenInference and Cloudflare is 12× for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Intelligence index
- 34.5
- Coding index
- 69.1
- Agentic index
- 41.7
About DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
Specifications
| Model ID | deepseek/deepseek-v4-flash-0731 |
|---|---|
| Provider | DeepSeek |
| Context window | 1.3M tokens |
| Max output | 944K tokens |
| Input modalities | text |
| Output modalities | text |
| Knowledge cutoff | - |
| Open weights | Yes - deepseek-ai/DeepSeek-V4-Flash-0731 |
| Released | July 31, 2026 |
Cheaper alternatives
Models that cost less than DeepSeek V4 Flash 0731 while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does DeepSeek: DeepSeek V4 Flash 0731 cost?
$0.060 per million input tokens and $0.120 per million output tokens. Cached input reads cost $0.012 per million tokens.
What is the context window of DeepSeek: DeepSeek V4 Flash 0731?
1.3M tokens, with up to 944K tokens of output per request.
Which provider serves DeepSeek: DeepSeek V4 Flash 0731 cheapest?
OpenInference at $0.040 per million input tokens - 12× cheaper than Cloudflare, the most expensive of the 27 hosts serving it.
Confirm against the source: DeepSeek official pricing.