Llama 4 Maverick
Meta's 400B mixture of experts of April 2025, with 17B active per token and one routed expert picked out of 128.
- Total
- 400B
- Active
- 17B
- Experts
- 128 routed + 1 shared1 per token
- Attention
- iRoPE, interleaved NoPE layers
- Context
- 1M tokens
- Input
- text, image
- Output
- text
- Licence
- Llama 4 Community License
- Weights
- Hugging Face
Contents1 to 3
How it is built
A Llama 4 Maverick token passes through 17B of the model's 400B parameters, about 4.3%.11Meta, Llama 4 Maverick model card, Hugging Face, 5 April 2025. https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct. Arithmetic: 17 / 400 = 4.25%. Dense layers alternate with mixture of experts layers.22Meta, "The Llama 4 herd", 5 April 2025. https://ai.meta.com/blog/llama-4-multimodal-intelligence/ Each MoE layer holds 128 routed experts and one shared expert. A token goes to the shared expert and to a single routed one.22
Meta calls the attention design iRoPE. Layers without positional embeddings are interleaved with layers that use RoPE.22 At inference time Meta also scales the attention temperature, to help the model generalise to longer inputs.22
Text and image tokens enter one backbone from the start. Meta calls this early fusion.22
What is new
These were Meta's first mixture-of-experts models. Maverick and Llama 4 Scout came out together on 5 April 2025.22
The pre-training mix held more than 30 trillion tokens, more than double the Llama 3 mix.22 Maverick itself saw about 22 trillion.11 Training ran in FP8. Hyper-parameters such as per-layer learning rates were set by a method Meta calls MetaP.22
Maverick was codistilled from Llama 4 Behemoth, a larger teacher that was still training when Maverick shipped.22
The card puts the instruction-tuned model at 80.5 on MMLU Pro, 0-shot. GPQA Diamond gives 69.8.11 Its knowledge stops at August 2024.11
Running it
Meta ships Maverick in BF16 and in FP8, behind a gate on Hugging Face.11 It "can be run on a single NVIDIA H100 DGX host", Meta writes.22 In FP8 the weights alone come to about 400 GB (400B x 1 byte).33Arithmetic on the model card figures: 400B parameters x 1 byte in FP8 = 400 GB, before the KV cache. https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct
The Llama 4 Community License sets two conditions. A licensee with more than 700 million monthly active users on the release date must ask Meta for a licence. Anyone who distributes the model must display "Built with Llama".44Meta, Llama 4 Community License Agreement, effective 5 April 2025. https://github.com/meta-llama/llama-models/blob/main/models/llama4/LICENSE
Under the acceptable use policy, individuals domiciled in the EU get no rights to Llama 4's multimodal models. Neither do companies with their principal place of business there. End users of a product built on these models are exempt.55Meta, Llama 4 Acceptable Use Policy. https://github.com/meta-llama/llama-models/blob/main/models/llama4/USE_POLICY.md
The newest repository in Meta's meta-llama organisation on Hugging Face dates from April 2025.66Hugging Face, meta-llama organisation sorted by creation date, read 29 September 2026. The newest repositories are Llama Guard 4 and Llama Prompt Guard 2, both of 29 April 2025. https://huggingface.co/meta-llama llama.com now redirects to Meta's developer site, which leads with the Muse models.77llama.com answers with a 301 redirect to developer.meta.com/ai, which redirects to dev.meta.ai, read 29 September 2026. https://dev.meta.ai/ The Meta Model API lists no Llama model at all.88Meta, Meta Model API documentation, models list, read 29 September 2026. https://dev.meta.ai/docs/models Meta's latest open-weight release is Muse Glimmer 30B, from August 2026.99Meta for Developers, "Build with Muse Glimmer", 12 August 2026. https://dev.meta.ai/resources/blog/build-with-muse-glimmer
Notes
-
Meta, Llama 4 Maverick model card, Hugging Face, 5 April 2025. https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct. Arithmetic: 17 / 400 = 4.25%. 2 3 4 5
-
Meta, "The Llama 4 herd", 5 April 2025. https://ai.meta.com/blog/llama-4-multimodal-intelligence/ 2 3 4 5 6 7 8 9 10
-
Arithmetic on the model card figures: 400B parameters x 1 byte in FP8 = 400 GB, before the KV cache. https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct
-
Meta, Llama 4 Community License Agreement, effective 5 April 2025. https://github.com/meta-llama/llama-models/blob/main/models/llama4/LICENSE
-
Meta, Llama 4 Acceptable Use Policy. https://github.com/meta-llama/llama-models/blob/main/models/llama4/USE_POLICY.md
-
Hugging Face, meta-llama organisation sorted by creation date, read 29 September 2026. The newest repositories are Llama Guard 4 and Llama Prompt Guard 2, both of 29 April 2025. https://huggingface.co/meta-llama
-
llama.com answers with a 301 redirect to developer.meta.com/ai, which redirects to dev.meta.ai, read 29 September 2026. https://dev.meta.ai/
-
Meta, Meta Model API documentation, models list, read 29 September 2026. https://dev.meta.ai/docs/models
-
Meta for Developers, "Build with Muse Glimmer", 12 August 2026. https://dev.meta.ai/resources/blog/build-with-muse-glimmer
More from Meta
- Muse Spark 1.3TotalUndisclosedActiveUndisclosed
- Muse Glimmer 30BTotal29.6BActive29.6B
- Llama 4 ScoutTotal109BActive17B