Google: Gemma 4 31B
by GoogleGoogle: Gemma 4 31B is a large language model from Google. It costs $0.090 per million input tokens and $0.340 per million output tokens. Its context window is 262K tokens.
- Input / 1M tokens
- $0.090
- Output / 1M tokens
- $0.340
- Cached input / 1M
- $0.050
- Context window
- 262K
On repeated prefixes
tokens
Who serves it cheapest
12 hosts serve Gemma 4 31B. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| DeepInfraCheapestfp4 | $0.090 | $0.340 | 262K | - | 99.4% |
| CoreWeavefp4 | $0.100 | $0.340 | 262K | - | 97.4% |
| Venicebf16 | $0.120 | $0.360 | 256K | - | 99.3% |
| Chutesfp4 | $0.120 | $0.370 | 131K | - | 91.7% |
| Crusoe | $0.140 | $0.400 | 262K | - | 96.8% |
| Friendli | $0.140 | $0.400 | 262K | - | 98.7% |
| Novitabf16 | $0.140 | $0.400 | 262K | - | 82.3% |
| Parasailfp8 | $0.150 | $0.400 | 262K | - | 97.5% |
| Together | $0.390 | $0.970 | 262K | - | 93.4% |
| SambaNova | $0.380 | $1.15 | 131K | - | 93.6% |
| ModelRunfp4 | $0.750 | $1.00 | 262K | - | 99.9% |
| SiliconFlowfp8 | $0.750 | $1.00 | 262K | - | 71.3% |
The spread between DeepInfra and SiliconFlow is 5.3× for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Intelligence index
- 15.4
- Coding index
- 43.4
- Agentic index
- 6.7
About Gemma 4 31B
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Specifications
| Model ID | google/gemma-4-31b-it |
|---|---|
| Provider | |
| Context window | 262K tokens |
| Max output | 16K tokens |
| Input modalities | image, text, video |
| Output modalities | text |
| Knowledge cutoff | - |
| Open weights | Yes - google/gemma-4-31B-it |
| Released | April 2, 2026 |
Cheaper alternatives
Models that cost less than Gemma 4 31B while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Google: Gemma 4 31B cost?
$0.090 per million input tokens and $0.340 per million output tokens. Cached input reads cost $0.050 per million tokens.
What is the context window of Google: Gemma 4 31B?
262K tokens, with up to 16K tokens of output per request.
Which provider serves Google: Gemma 4 31B cheapest?
DeepInfra at $0.090 per million input tokens - 5.3× cheaper than SiliconFlow, the most expensive of the 12 hosts serving it.
Confirm against the source: Google official pricing.