Salta la copertina

MetaIn inglese

Llama 4 Scout

A 109B Llama 4 model with 16 experts that reads up to 10 million tokens and fits on one H100 at 4 bits.

Totale
109B
Attivi
17B
Esperti
16 instradati + 1 condiviso1 per token
Attenzione
iRoPE, interleaved NoPE layers
Contesto
10M token
Ingresso
text, image
Uscita
text
Licenza
Llama 4 Community License

Indiceda 1 a 3

How it is built

Llama 4 Scout has 16 experts and 109B parameters. Each token uses 17B of them, about 15.6%.11Meta, Llama 4 Scout model card, Hugging Face, 5 April 2025. MMLU Pro 74.3, GPQA Diamond 57.2, LiveCodeBench 32.8 pass@1, DocVQA 94.4 ANLS, all 0-shot. https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct. Arithmetic: 17 / 109 = 15.6%.

The card does not say how many experts a token visits, or whether one of them is shared.11 Meta's reference code for Llama 4 builds a shared expert into every mixture of experts layer and sends each token to one routed expert by default.22Meta, Llama 4 reference implementation in llama-models: moe.py creates a shared_expert in every MoE layer, and args.py sets top_k to 1 by default. Scout's own params file sits behind the Hugging Face gate. https://github.com/meta-llama/llama-models/tree/main/models/llama4 Scout's own configuration sits behind the Hugging Face gate.22

NoPE layers
Attention layers that add no positional embedding to their queries and keys.

Attention is iRoPE, as in Llama 4 Maverick. Some layers carry no positional embeddings, and the attention temperature is scaled at inference time.33Meta, "The Llama 4 herd", 5 April 2025: Scout "is both pre-trained and post-trained with a 256K context length, which empowers the base model with advanced length generalization capability". https://ai.meta.com/blog/llama-4-multimodal-intelligence/

NoPE layers
Attention layers that add no positional embedding to their queries and keys.

Images and text go through the same backbone, a design Meta calls early fusion.33

What is new

Scout reads up to 10M tokens of context, ten times Maverick's 1M.1144Meta, Llama 4 Maverick model card, Hugging Face. https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct. Arithmetic: 10M / 1M = 10; 40 / 22 = 1.8. Meta pre-trained and post-trained it at 256K, and credits that training for its length generalisation.33

Pre-training used about 40 trillion tokens, nearly twice Maverick's 22 trillion.1144 Training took 5.0 million H100 GPU hours. Meta puts the location-based emissions at 1,354 tons of CO2 equivalent.11

Instruction-tuned, Scout scores 74.3 on MMLU Pro and 57.2 on GPQA Diamond.11 LiveCodeBench, over problems from October 2024 to February 2025, gives 32.8 pass@1. On DocVQA its 94.4 ANLS equals Maverick's.1144 All four scores are 0-shot.11 The knowledge cutoff is August 2024.11

Running it

The weights are released in BF16.55Meta, Llama 4 model card in llama-models: "The Llama 4 Scout model is released as BF16 weights". https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md At 2 bytes per parameter that is about 218 GB (109B x 2 bytes).66Arithmetic, weights only: 109B x 2 bytes = 218 GB in BF16; 109B x 0.5 byte = 54.5 GB at 4 bits, before quantization scales and the KV cache. The card gives training hardware as H100-80GB. https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct

The card says Scout "can fit within a single H100 GPU with on-the-fly int4 quantization".11 At half a byte per parameter the weights shrink to about 55 GB (109B x 0.5 byte), which leaves part of an 80 GB H100 for the KV cache.66

Access to the weights on Hugging Face is gated. Companies above 700 million monthly active users need a separate licence from Meta.77Meta, Llama 4 Community License Agreement, effective 5 April 2025. https://github.com/meta-llama/llama-models/blob/main/models/llama4/LICENSE The acceptable use policy grants no rights to Llama 4's multimodal models to people domiciled in the EU, or to companies based there. End users of products built on them are exempt.88Meta, Llama 4 Acceptable Use Policy. https://github.com/meta-llama/llama-models/blob/main/models/llama4/USE_POLICY.md

No Llama model appears in the Meta Model API.99Meta, Meta Model API documentation, models list, read 29 September 2026. https://dev.meta.ai/docs/models Meta's newest open weights belong to Muse Glimmer 30B, a dense model from August 2026.1010Meta Superintelligence Labs, Muse Glimmer model card, August 2026. https://huggingface.co/meta-models/Muse-Glimmer-30B

Note

  1. Meta, Llama 4 Scout model card, Hugging Face, 5 April 2025. MMLU Pro 74.3, GPQA Diamond 57.2, LiveCodeBench 32.8 pass@1, DocVQA 94.4 ANLS, all 0-shot. https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct. Arithmetic: 17 / 109 = 15.6%. 2 3 4 5 6 7 8 9 10

  2. Meta, Llama 4 reference implementation in llama-models: moe.py creates a shared_expert in every MoE layer, and args.py sets top_k to 1 by default. Scout's own params file sits behind the Hugging Face gate. https://github.com/meta-llama/llama-models/tree/main/models/llama4 2

  3. Meta, "The Llama 4 herd", 5 April 2025: Scout "is both pre-trained and post-trained with a 256K context length, which empowers the base model with advanced length generalization capability". https://ai.meta.com/blog/llama-4-multimodal-intelligence/ 2 3

  4. Meta, Llama 4 Maverick model card, Hugging Face. https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct. Arithmetic: 10M / 1M = 10; 40 / 22 = 1.8. 2 3

  5. Meta, Llama 4 model card in llama-models: "The Llama 4 Scout model is released as BF16 weights". https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md

  6. Arithmetic, weights only: 109B x 2 bytes = 218 GB in BF16; 109B x 0.5 byte = 54.5 GB at 4 bits, before quantization scales and the KV cache. The card gives training hardware as H100-80GB. https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct 2

  7. Meta, Llama 4 Community License Agreement, effective 5 April 2025. https://github.com/meta-llama/llama-models/blob/main/models/llama4/LICENSE

  8. Meta, Llama 4 Acceptable Use Policy. https://github.com/meta-llama/llama-models/blob/main/models/llama4/USE_POLICY.md

  9. Meta, Meta Model API documentation, models list, read 29 September 2026. https://dev.meta.ai/docs/models

  10. Meta Superintelligence Labs, Muse Glimmer model card, August 2026. https://huggingface.co/meta-models/Muse-Glimmer-30B

Altri modelli di Meta

Tutti i modelli di Meta