Salta la copertina

Google DeepMindIn inglese

Gemma 4 26B A4B

A mixture of experts that runs 3.8B of its 25.2B parameters per token, built by Google for latency, with open weights under Apache 2.0.

Totale
25,2B
Attivi
3,8B
Esperti
128 instradati + 1 condiviso8 per token
Strati
30
Attenzione
Hybrid: sliding window and global
Contesto
256k token
Ingresso
text, image
Uscita
text
Licenza
Apache 2.0

Indiceda 1 a 3

How it is built

Of the 25.2B parameters in Gemma 4 26B A4B, 3.8B work on any one token.11Google, Gemma 4 model card: 26B A4B, 25.2B total parameters, 3.8B active, 30 layers, experts "8 active / 128 total and 1 shared", sliding window 1,024 tokens, context 256K. https://ai.google.dev/gemma/docs/core/model_card_4 That is about 15% of the weights (3.8 / 25.2).22Arithmetic on the card's figures: 3.8 / 25.2 = 0.151. https://ai.google.dev/gemma/docs/core/model_card_4

Shared expert
An expert that every token passes through. The router chooses only among the others.

The model is a mixture of experts with 30 layers, each holding 128 experts.11 The router picks 8 of them for every token, and 1 shared expert runs on all tokens.11 The config gives each expert an inner width of 704.33Google, config.json of gemma-4-26B-A4B-it on Hugging Face: enable_moe_block true for the whole model, num_experts 128, top_k_experts 8, moe_intermediate_size 704, 30 layers with full attention at layers 6, 12, 18, 24 and 30, 16 attention heads, 8 key-value heads, 2 global key-value heads, dtype bfloat16. https://huggingface.co/google/gemma-4-26B-A4B-it/blob/main/config.json. Arithmetic: 30 / 6 = 5 global layers; 30 - 5 = 25 local layers.

Shared expert
An expert that every token passes through. The router chooses only among the others.

Attention is laid out as in the dense Gemma 4 31B. Five layers in six look back over a sliding window of 1,024 tokens, and the sixth sees the whole context.1133 Across 30 layers that makes 25 local and 5 global.33

The global layers keep 2 key and value heads for 16 query heads. The local ones keep 8.33

Fig. 1
Token

30 layers. In each one, a token goes to 8 of 128 routed experts and to 1 shared expert. The experts it meets here are simulated, because the real router depends on the trained weights.

What is new

It came out in April 2026 with the dense 31B and the small E2B and E4B models.44Google, Gemma releases page: "Release of Gemma 4" in E2B, E4B, 31B and 26B A4B sizes, dated 31 March 2026; Gemma 4 MTP drafters for the same sizes, 16 April 2026. Google's launch post and the Gemini API release notes date the launch 2 April 2026, the day of the first commit to the Hugging Face repository. https://ai.google.dev/gemma/docs/releases

Google's launch post frames it as the latency model. It activates "only 3.8 billion of its total parameters during inference".55Google, "Gemma 4: Byte for byte, the most capable open models", 2 April 2026: the 26B MoE has a "focus on latency, activating only 3.8 billion of its total parameters during inference"; unquantized weights "fit efficiently on a single 80GB NVIDIA H100 GPU"; the 26B holds "the #6 spot" among open models on the Arena AI text leaderboard at launch. https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ At launch it held sixth place among open models on the Arena AI text leaderboard.55

On the Hugging Face card the gap to the 31B is small in knowledge and maths. MMLU Pro gives 82.6% against 85.2%, and AIME 2026 without tools 88.3% against 89.2%.66Google, gemma-4-26B-A4B-it model card on Hugging Face: licence Apache 2.0, not gated. Instruction-tuned results, 26B A4B against 31B: MMLU Pro 82.6% and 85.2%; AIME 2026 no tools 88.3% and 89.2%; Codeforces ELO 1718 and 2150; MRCR v2 8 needle 128k 44.1% and 66.4%. Thinking is enabled by a token at the start of the system prompt. https://huggingface.co/google/gemma-4-26B-A4B-it

The gap widens on code and long context. The Codeforces Elo is 1718, where the dense model reaches 2150.66 MRCR v2 with eight needles at 128k tokens gives 44.1%, against 66.4%.66

Running it

All 128 experts have to be loaded, whichever ones the router picks. Google's memory table puts the weights at 57.7 GB in BF16, only 12.2 GB less than the dense 31B.77Google, Gemma 4 overview, memory table: 26B A4B at 57.7 GB in BF16, 28.8 GB in SFP8 and 14.4 GB in Q4_0; 31B at 69.9 GB in BF16. The figures "only account for the memory required to load the static model weights". https://ai.google.dev/gemma/docs/core. Arithmetic: 69.9 - 57.7 = 12.2 GB. In SFP8 they take 28.8 GB, in Q4_0 14.4 GB.77 The table counts the weights and nothing else.77

The launch post says the unquantized weights fit on one 80GB NVIDIA H100.55 The licence is Apache 2.0, and the repository has no access gate.66

Speculative decoding
A small draft model proposes the next few tokens, and the large model checks them in one pass.

A <|think|> token at the start of the system prompt turns thinking on.66 Google added a draft model for speculative decoding on 16 April 2026.44

Speculative decoding
A small draft model proposes the next few tokens, and the large model checks them in one pass.

The Gemini API serves it too, as gemma-4-26b-a4b-it.88Google, Gemini API release notes, 2 April 2026: "Released gemma-4-26b-a4b-it and gemma-4-31b-it". https://ai.google.dev/gemini-api/docs/changelog

Note

  1. Google, Gemma 4 model card: 26B A4B, 25.2B total parameters, 3.8B active, 30 layers, experts "8 active / 128 total and 1 shared", sliding window 1,024 tokens, context 256K. https://ai.google.dev/gemma/docs/core/model_card_4 2 3 4

  2. Arithmetic on the card's figures: 3.8 / 25.2 = 0.151. https://ai.google.dev/gemma/docs/core/model_card_4

  3. Google, config.json of gemma-4-26B-A4B-it on Hugging Face: enable_moe_block true for the whole model, num_experts 128, top_k_experts 8, moe_intermediate_size 704, 30 layers with full attention at layers 6, 12, 18, 24 and 30, 16 attention heads, 8 key-value heads, 2 global key-value heads, dtype bfloat16. https://huggingface.co/google/gemma-4-26B-A4B-it/blob/main/config.json. Arithmetic: 30 / 6 = 5 global layers; 30 - 5 = 25 local layers. 2 3 4

  4. Google, Gemma releases page: "Release of Gemma 4" in E2B, E4B, 31B and 26B A4B sizes, dated 31 March 2026; Gemma 4 MTP drafters for the same sizes, 16 April 2026. Google's launch post and the Gemini API release notes date the launch 2 April 2026, the day of the first commit to the Hugging Face repository. https://ai.google.dev/gemma/docs/releases 2

  5. Google, "Gemma 4: Byte for byte, the most capable open models", 2 April 2026: the 26B MoE has a "focus on latency, activating only 3.8 billion of its total parameters during inference"; unquantized weights "fit efficiently on a single 80GB NVIDIA H100 GPU"; the 26B holds "the #6 spot" among open models on the Arena AI text leaderboard at launch. https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ 2 3

  6. Google, gemma-4-26B-A4B-it model card on Hugging Face: licence Apache 2.0, not gated. Instruction-tuned results, 26B A4B against 31B: MMLU Pro 82.6% and 85.2%; AIME 2026 no tools 88.3% and 89.2%; Codeforces ELO 1718 and 2150; MRCR v2 8 needle 128k 44.1% and 66.4%. Thinking is enabled by a token at the start of the system prompt. https://huggingface.co/google/gemma-4-26B-A4B-it 2 3 4 5

  7. Google, Gemma 4 overview, memory table: 26B A4B at 57.7 GB in BF16, 28.8 GB in SFP8 and 14.4 GB in Q4_0; 31B at 69.9 GB in BF16. The figures "only account for the memory required to load the static model weights". https://ai.google.dev/gemma/docs/core. Arithmetic: 69.9 - 57.7 = 12.2 GB. 2 3

  8. Google, Gemini API release notes, 2 April 2026: "Released gemma-4-26b-a4b-it and gemma-4-31b-it". https://ai.google.dev/gemini-api/docs/changelog

Altri modelli di Google DeepMind

Tutti i modelli di Google DeepMind