Salta la copertina

Google DeepMindIn inglese

Gemma 4 31B

The largest Gemma 4 model: dense, 30.7B parameters, open weights under Apache 2.0 and a context window of 256K tokens.

Totale
30,7B
Attivi
30,7B
Esperti
Nessunomodello denso
Strati
60
Attenzione
Hybrid: sliding window and global
Contesto
256k token
Ingresso
text, image
Uscita
text
Licenza
Apache 2.0

Indiceda 1 a 3

How it is built

Every token goes through all 30.7B parameters of Gemma 4 31B, spread over 60 layers.11Google, Gemma 4 model card: 31B dense, 30.7B parameters, 60 layers, sliding window 1,024 tokens, context 256K, vocabulary 262K, vision encoder about 550M parameters; hybrid attention "ensuring the final layer is always global". https://ai.google.dev/gemma/docs/core/model_card_4

Sliding window attention
Each token looks back over a fixed number of earlier tokens. The memory these layers keep stops growing once the window is full.

Most layers see only the last 1,024 tokens, through a sliding window. The rest attend to the whole context.11

Sliding window attention
Each token looks back over a fixed number of earlier tokens. The memory these layers keep stops growing once the window is full.

The published config makes every sixth layer global, 10 in all, with 50 local layers between them.22Google, config.json of gemma-4-31B-it on Hugging Face: 60 layers, full attention at layers 6, 12, 18 and so on up to 60, 32 attention heads, 16 key-value heads, 4 global key-value heads, dtype bfloat16. https://huggingface.co/google/gemma-4-31B-it/blob/main/config.json. Arithmetic: 60 / 6 = 10 global layers; 60 - 10 = 50 local layers. The card adds that the final layer is always global.11

The local layers pair 32 query heads with 16 key and value heads. The global layers cut that to 4.22 Images go through a vision encoder of about 550M parameters, and the vocabulary has 262K entries.11

Fig. 1

60 layers, all dense.

What is new

Gemma 4 came out in April 2026 in four sizes.33Google, Gemma releases page: "Release of Gemma 4" in E2B, E4B, 31B and 26B A4B sizes, dated 31 March 2026; Gemma 4 MTP drafters for the same sizes, 16 April 2026. Google's launch post and the Gemini API release notes date the launch 2 April 2026, the day of the first commit to the Hugging Face repository. https://ai.google.dev/gemma/docs/releases

The licence changed. Gemma 3 27B was released under Google's own Gemma licence, with gated weights on Hugging Face.44Hugging Face, google/gemma-3-27b-it repository: licence "gemma", access gated. https://huggingface.co/google/gemma-3-27b-it Gemma 4 uses Apache 2.0, and nobody has to ask for access.55Google, gemma-4-31B-it model card on Hugging Face: licence Apache 2.0, not gated. Instruction-tuned results: MMLU Pro 85.2% against 67.6% for Gemma 3 27B; AIME 2026 no tools 89.2% against 20.8%; LiveCodeBench v6 80.0% against 29.1%; Codeforces ELO 2150 against 110; the Gemma 3 27B column is headed "no think". Thinking is enabled by a token at the start of the system prompt. https://huggingface.co/google/gemma-4-31B-it

The context window doubled, from 128K tokens in Gemma 3 27B to 256K.66Google, Gemma 3 model card: "Total input context of 128K tokens for the 4B, 12B, and 27B sizes". https://ai.google.dev/gemma/docs/core/model_card_3. Arithmetic: 256K / 128K = 2. 11

Against Gemma 3 27B the benchmark gaps are wide, though the older model's column was run without thinking.55 MMLU Pro rose from 67.6% to 85.2%. AIME 2026 without tools went from 20.8% to 89.2%, and LiveCodeBench v6 from 29.1% to 80.0%.55 The older model's Codeforces Elo is 110. Gemma 4 31B reaches 2150.55

Google's launch post put it third among open models on the Arena AI text leaderboard.77Google, "Gemma 4: Byte for byte, the most capable open models", 2 April 2026: the 31B ranked "#3 open model in the world" on the Arena AI text leaderboard at launch; weights on Hugging Face, Kaggle and Ollama. https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/

Running it

The weights are stored in BF16, where Google's memory table puts them at 69.9 GB.2288Google, Gemma 4 overview, memory table: 31B at 69.9 GB in BF16, 34.9 GB in SFP8 and 17.5 GB in Q4_0; 26B A4B at 57.7 GB in BF16. The figures "only account for the memory required to load the static model weights". https://ai.google.dev/gemma/docs/core. Arithmetic: 69.9 / 2 = 34.95. SFP8 halves that to 34.9 GB, and Q4_0 brings it to 17.5 GB.88 These figures cover the weights alone. The context and the serving software come on top.88

Gemma 4 26B A4B, the mixture of experts released with it, needs 57.7 GB in BF16.88

Thinking stays off unless the system prompt starts with the <|think|> token.55 Draft models for speculative decoding arrived on 16 April 2026.33

The weights are on Hugging Face and Kaggle, and the Gemini API serves the model as gemma-4-31b-it.7799Google, Gemini API release notes, 2 April 2026: "Released gemma-4-26b-a4b-it and gemma-4-31b-it". https://ai.google.dev/gemini-api/docs/changelog

Note

  1. Google, Gemma 4 model card: 31B dense, 30.7B parameters, 60 layers, sliding window 1,024 tokens, context 256K, vocabulary 262K, vision encoder about 550M parameters; hybrid attention "ensuring the final layer is always global". https://ai.google.dev/gemma/docs/core/model_card_4 2 3 4 5

  2. Google, config.json of gemma-4-31B-it on Hugging Face: 60 layers, full attention at layers 6, 12, 18 and so on up to 60, 32 attention heads, 16 key-value heads, 4 global key-value heads, dtype bfloat16. https://huggingface.co/google/gemma-4-31B-it/blob/main/config.json. Arithmetic: 60 / 6 = 10 global layers; 60 - 10 = 50 local layers. 2 3

  3. Google, Gemma releases page: "Release of Gemma 4" in E2B, E4B, 31B and 26B A4B sizes, dated 31 March 2026; Gemma 4 MTP drafters for the same sizes, 16 April 2026. Google's launch post and the Gemini API release notes date the launch 2 April 2026, the day of the first commit to the Hugging Face repository. https://ai.google.dev/gemma/docs/releases 2

  4. Hugging Face, google/gemma-3-27b-it repository: licence "gemma", access gated. https://huggingface.co/google/gemma-3-27b-it

  5. Google, gemma-4-31B-it model card on Hugging Face: licence Apache 2.0, not gated. Instruction-tuned results: MMLU Pro 85.2% against 67.6% for Gemma 3 27B; AIME 2026 no tools 89.2% against 20.8%; LiveCodeBench v6 80.0% against 29.1%; Codeforces ELO 2150 against 110; the Gemma 3 27B column is headed "no think". Thinking is enabled by a token at the start of the system prompt. https://huggingface.co/google/gemma-4-31B-it 2 3 4 5

  6. Google, Gemma 3 model card: "Total input context of 128K tokens for the 4B, 12B, and 27B sizes". https://ai.google.dev/gemma/docs/core/model_card_3. Arithmetic: 256K / 128K = 2.

  7. Google, "Gemma 4: Byte for byte, the most capable open models", 2 April 2026: the 31B ranked "#3 open model in the world" on the Arena AI text leaderboard at launch; weights on Hugging Face, Kaggle and Ollama. https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ 2

  8. Google, Gemma 4 overview, memory table: 31B at 69.9 GB in BF16, 34.9 GB in SFP8 and 17.5 GB in Q4_0; 26B A4B at 57.7 GB in BF16. The figures "only account for the memory required to load the static model weights". https://ai.google.dev/gemma/docs/core. Arithmetic: 69.9 / 2 = 34.95. 2 3 4

  9. Google, Gemini API release notes, 2 April 2026: "Released gemma-4-26b-a4b-it and gemma-4-31b-it". https://ai.google.dev/gemini-api/docs/changelog

Altri modelli di Google DeepMind

Tutti i modelli di Google DeepMind