Skip to content
LLMs
Google logo

Google: Gemma 4 31B

by Google

Google: Gemma 4 31B is a large language model from Google. It costs $0.090 per million input tokens and $0.340 per million output tokens. Its context window is 262K tokens.

Free tier availableReasoningTool callingStructured outputPrompt cachingBatch tierOpen weightsimage inputvideo input
Input / 1M tokens
$0.090
Output / 1M tokens
$0.340
Cached input / 1M
$0.050

On repeated prefixes

Context window
262K

tokens

Who serves it cheapest

12 hosts serve Gemma 4 31B. Same weights, same API - the price difference is pure margin and routing.

Providers serving Google: Gemma 4 31B, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
DeepInfraCheapestfp4$0.090$0.340262K-99.4%
CoreWeavefp4$0.100$0.340262K-97.4%
Venicebf16$0.120$0.360256K-99.3%
Chutesfp4$0.120$0.370131K-91.7%
Crusoe$0.140$0.400262K-96.8%
Friendli$0.140$0.400262K-98.7%
Novitabf16$0.140$0.400262K-82.3%
Parasailfp8$0.150$0.400262K-97.5%
Together$0.390$0.970262K-93.4%
SambaNova$0.380$1.15131K-93.6%
ModelRunfp4$0.750$1.00262K-99.9%
SiliconFlowfp8$0.750$1.00262K-71.3%

The spread between DeepInfra and SiliconFlow is 5.3× for identical weights. Quantization and context limits differ, so check both columns before switching.

Benchmarks

Independent scores published alongside the catalogue.

Intelligence index
15.4
Coding index
43.4
Agentic index
6.7

About Gemma 4 31B

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Specifications

Google: Gemma 4 31B specifications
Model IDgoogle/gemma-4-31b-it
ProviderGoogle
Context window262K tokens
Max output16K tokens
Input modalitiesimage, text, video
Output modalitiestext
Knowledge cutoff-
Open weightsYes - google/gemma-4-31B-it
ReleasedApril 2, 2026

Cheaper alternatives

Models that cost less than Gemma 4 31B while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does Google: Gemma 4 31B cost?

$0.090 per million input tokens and $0.340 per million output tokens. Cached input reads cost $0.050 per million tokens.

What is the context window of Google: Gemma 4 31B?

262K tokens, with up to 16K tokens of output per request.

Which provider serves Google: Gemma 4 31B cheapest?

DeepInfra at $0.090 per million input tokens - 5.3× cheaper than SiliconFlow, the most expensive of the 12 hosts serving it.

Confirm against the source: Google official pricing.