Google: Gemma 3 4B
by GoogleGoogle: Gemma 3 4B is a large language model from Google. It costs $0.050 per million input tokens and $0.100 per million output tokens. Its context window is 131K tokens.
- Input / 1M tokens
- $0.050
- Output / 1M tokens
- $0.100
- Cached input / 1M
- -
- Context window
- 131K
Not supported
tokens
Who serves it cheapest
1 hosts serve Gemma 3 4B. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| DeepInfraCheapestbf16 | $0.050 | $0.100 | 131K | - | 100.0% |
Benchmarks
Independent scores published alongside the catalogue.
- Coding index
- 2.7
About Gemma 3 4B
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
Specifications
| Model ID | google/gemma-3-4b-it |
|---|---|
| Provider | |
| Context window | 131K tokens |
| Max output | 16K tokens |
| Input modalities | text, image |
| Output modalities | text |
| Knowledge cutoff | 2024-08-31 |
| Open weights | Yes - google/gemma-3-4b-it |
| Released | March 13, 2025 |
Cheaper alternatives
Models that cost less than Gemma 3 4B while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Google: Gemma 3 4B cost?
$0.050 per million input tokens and $0.100 per million output tokens.
What is the context window of Google: Gemma 3 4B?
131K tokens, with up to 16K tokens of output per request.
Confirm against the source: Google official pricing.