Gemma 4 31B
The largest Gemma 4 model: dense, 30.7B parameters, open weights under Apache 2.0 and a context window of 256K tokens.
- Total
- 30.7B
- Active
- 30.7B
- Experts
- Nonedense model
- Layers
- 60
- Attention
- Hybrid: sliding window and global
- Context
- 256k tokens
- Input
- text, image
- Output
- text
- Licence
- Apache 2.0
- Weights
- Hugging Face
Contents1 to 3
How it is built
Every token goes through all 30.7B parameters of Gemma 4 31B, spread over 60 layers.11Google, Gemma 4 model card: 31B dense, 30.7B parameters, 60 layers, sliding window 1,024 tokens, context 256K, vocabulary 262K, vision encoder about 550M parameters; hybrid attention "ensuring the final layer is always global". https://ai.google.dev/gemma/docs/core/model_card_4
Most layers see only the last 1,024 tokens, through a sliding window. The rest attend to the whole context.11
The published config makes every sixth layer global, 10 in all, with 50 local layers between them.22Google, config.json of gemma-4-31B-it on Hugging Face: 60 layers, full attention at layers 6, 12, 18 and so on up to 60, 32 attention heads, 16 key-value heads, 4 global key-value heads, dtype bfloat16. https://huggingface.co/google/gemma-4-31B-it/blob/main/config.json. Arithmetic: 60 / 6 = 10 global layers; 60 - 10 = 50 local layers. The card adds that the final layer is always global.11
The local layers pair 32 query heads with 16 key and value heads. The global layers cut that to 4.22 Images go through a vision encoder of about 550M parameters, and the vocabulary has 262K entries.11
What is new
Gemma 4 came out in April 2026 in four sizes.33Google, Gemma releases page: "Release of Gemma 4" in E2B, E4B, 31B and 26B A4B sizes, dated 31 March 2026; Gemma 4 MTP drafters for the same sizes, 16 April 2026. Google's launch post and the Gemini API release notes date the launch 2 April 2026, the day of the first commit to the Hugging Face repository. https://ai.google.dev/gemma/docs/releases
The licence changed. Gemma 3 27B was released under Google's own Gemma licence, with gated weights on Hugging Face.44Hugging Face, google/gemma-3-27b-it repository: licence "gemma", access gated. https://huggingface.co/google/gemma-3-27b-it Gemma 4 uses Apache 2.0, and nobody has to ask for access.55Google, gemma-4-31B-it model card on Hugging Face: licence Apache 2.0, not gated. Instruction-tuned results: MMLU Pro 85.2% against 67.6% for Gemma 3 27B; AIME 2026 no tools 89.2% against 20.8%; LiveCodeBench v6 80.0% against 29.1%; Codeforces ELO 2150 against 110; the Gemma 3 27B column is headed "no think". Thinking is enabled by a token at the start of the system prompt. https://huggingface.co/google/gemma-4-31B-it
The context window doubled, from 128K tokens in Gemma 3 27B to 256K.66Google, Gemma 3 model card: "Total input context of 128K tokens for the 4B, 12B, and 27B sizes". https://ai.google.dev/gemma/docs/core/model_card_3. Arithmetic: 256K / 128K = 2. 11
Against Gemma 3 27B the benchmark gaps are wide, though the older model's column was run without thinking.55 MMLU Pro rose from 67.6% to 85.2%. AIME 2026 without tools went from 20.8% to 89.2%, and LiveCodeBench v6 from 29.1% to 80.0%.55 The older model's Codeforces Elo is 110. Gemma 4 31B reaches 2150.55
Google's launch post put it third among open models on the Arena AI text leaderboard.77Google, "Gemma 4: Byte for byte, the most capable open models", 2 April 2026: the 31B ranked "#3 open model in the world" on the Arena AI text leaderboard at launch; weights on Hugging Face, Kaggle and Ollama. https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/
Running it
The weights are stored in BF16, where Google's memory table puts them at 69.9 GB.2288Google, Gemma 4 overview, memory table: 31B at 69.9 GB in BF16, 34.9 GB in SFP8 and 17.5 GB in Q4_0; 26B A4B at 57.7 GB in BF16. The figures "only account for the memory required to load the static model weights". https://ai.google.dev/gemma/docs/core. Arithmetic: 69.9 / 2 = 34.95. SFP8 halves that to 34.9 GB, and Q4_0 brings it to 17.5 GB.88 These figures cover the weights alone. The context and the serving software come on top.88
Gemma 4 26B A4B, the mixture of experts released with it, needs 57.7 GB in BF16.88
Thinking stays off unless the system prompt starts with the <|think|> token.55 Draft models for speculative decoding arrived on 16 April 2026.33
The weights are on Hugging Face and Kaggle, and the Gemini API serves the model as gemma-4-31b-it.7799Google, Gemini API release notes, 2 April 2026: "Released gemma-4-26b-a4b-it and gemma-4-31b-it". https://ai.google.dev/gemini-api/docs/changelog
Notes
-
Google, Gemma 4 model card: 31B dense, 30.7B parameters, 60 layers, sliding window 1,024 tokens, context 256K, vocabulary 262K, vision encoder about 550M parameters; hybrid attention "ensuring the final layer is always global". https://ai.google.dev/gemma/docs/core/model_card_4 2 3 4 5
-
Google, config.json of gemma-4-31B-it on Hugging Face: 60 layers, full attention at layers 6, 12, 18 and so on up to 60, 32 attention heads, 16 key-value heads, 4 global key-value heads, dtype bfloat16. https://huggingface.co/google/gemma-4-31B-it/blob/main/config.json. Arithmetic: 60 / 6 = 10 global layers; 60 - 10 = 50 local layers. 2 3
-
Google, Gemma releases page: "Release of Gemma 4" in E2B, E4B, 31B and 26B A4B sizes, dated 31 March 2026; Gemma 4 MTP drafters for the same sizes, 16 April 2026. Google's launch post and the Gemini API release notes date the launch 2 April 2026, the day of the first commit to the Hugging Face repository. https://ai.google.dev/gemma/docs/releases 2
-
Hugging Face, google/gemma-3-27b-it repository: licence "gemma", access gated. https://huggingface.co/google/gemma-3-27b-it
-
Google, gemma-4-31B-it model card on Hugging Face: licence Apache 2.0, not gated. Instruction-tuned results: MMLU Pro 85.2% against 67.6% for Gemma 3 27B; AIME 2026 no tools 89.2% against 20.8%; LiveCodeBench v6 80.0% against 29.1%; Codeforces ELO 2150 against 110; the Gemma 3 27B column is headed "no think". Thinking is enabled by a token at the start of the system prompt. https://huggingface.co/google/gemma-4-31B-it 2 3 4 5
-
Google, Gemma 3 model card: "Total input context of 128K tokens for the 4B, 12B, and 27B sizes". https://ai.google.dev/gemma/docs/core/model_card_3. Arithmetic: 256K / 128K = 2.
-
Google, "Gemma 4: Byte for byte, the most capable open models", 2 April 2026: the 31B ranked "#3 open model in the world" on the Arena AI text leaderboard at launch; weights on Hugging Face, Kaggle and Ollama. https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ 2
-
Google, Gemma 4 overview, memory table: 31B at 69.9 GB in BF16, 34.9 GB in SFP8 and 17.5 GB in Q4_0; 26B A4B at 57.7 GB in BF16. The figures "only account for the memory required to load the static model weights". https://ai.google.dev/gemma/docs/core. Arithmetic: 69.9 / 2 = 34.95. 2 3 4
-
Google, Gemini API release notes, 2 April 2026: "Released gemma-4-26b-a4b-it and gemma-4-31b-it". https://ai.google.dev/gemini-api/docs/changelog
More from Google DeepMind
- Gemini 3.1 ProTotalUndisclosedActiveUndisclosed
- Gemini 3.8 FlashTotalUndisclosedActiveUndisclosed
- Gemini 3.5 Flash-LiteTotalUndisclosedActiveUndisclosed
- Gemma 4 26B A4BTotal25.2BActive3.8B