Muse Glimmer 30B
Meta's open dense model of about 30B parameters, distilled from Muse Spark to run agents on a Mac or a single consumer GPU.
- Total
- 29.6B
- Active
- 29.6B
- Experts
- Nonedense model
- Layers
- 52
- Attention
- Local and global 3:1, gated GQA
- Context
- 131k tokens
- Input
- text, image
- Output
- text
- Licence
- Apache 2.0
- Weights
- Hugging Face
Contents1 to 3
How it is built
Muse Glimmer counts about 29.6B parameters, and that figure includes the vision encoder.11Meta Superintelligence Labs, Muse Glimmer model card, Hugging Face, August 2026. https://huggingface.co/meta-models/Muse-Glimmer-30B The encoder alone is a ViT-G/14 of about 1.8B. Meta's card calls the model "distilled from Muse Spark" and built "for autonomous agentic tasks on consumer hardware".11
It is a dense transformer. The language model has 52 layers with a hidden size of 6,656.11 Three local layers come before every global one, which gives 39 local and 13 global (52 x 3/4 and 52 x 1/4).22Meta, config.json in the official repository: num_hidden_layers 52, layer_types, sliding_window 2048, layer_rope_theta 0 on every full_attention layer, max_position_embeddings 131072. https://huggingface.co/meta-models/Muse-Glimmer-30B/blob/main/config.json. Arithmetic: 52 x 3/4 = 39 sliding layers and 52 x 1/4 = 13 full layers, as layer_types lists them. The local layers attend over a sliding window of 2,048 tokens.11
RoPE, with a base of 500,000, runs only in the local layers. In the config the global layers get no rotary position at all.22 Each layer has 32 query heads sharing 2 key-value heads, with gated attention.11 The feed-forward block is SwiGLU, 19,968 wide.11
What is new
The weights went public on Hugging Face on 10 August 2026, under Apache 2.0.33Hugging Face, launch post for Muse Glimmer, 10 August 2026, which says the model was "released today". This is the earliest official date; Meta's developer blog post is dated 12 August 2026. https://huggingface.co/blog/muse-glimmer Meta's developer blog calls it the most permissive licence Meta has used for an open model.44Meta for Developers, "Build with Muse Glimmer", 12 August 2026: "the most permissive license we've used for an open model"; "a default context window of 128K tokens". https://dev.meta.ai/resources/blog/build-with-muse-glimmer Llama 4 had come under the Llama 4 Community License.55Meta, Llama 4 Community License Agreement, effective 5 April 2025. https://github.com/meta-llama/llama-models/blob/main/models/llama4/LICENSE
For speculative decoding the model ships with a small DFlash drafter. It proposes 16 tokens in one forward pass, and the main model verifies them in parallel.11
Meta measured the drafter with its 17 GB quantization at batch size 1. On an NVIDIA RTX 5090 under llama.cpp, generation went from 74.9 to 233.4 tokens per second.11 An Apple M5 Max under ExecuTorch went from 26.6 to 50.2.11
The comparison on the card is with Gemma 4 31B and Qwen3.6-27B, both thinking. At high reasoning Muse Glimmer scores 75.5 on MCP Atlas (Public), where the other two score 54.2 and 62.5. On SWE-Bench Verified it trails Qwen, 76.0 against 77.2, with Gemma at 66.6.11
Running it
The repository has no access gate.66Hugging Face, repository metadata for meta-models/Muse-Glimmer-30B (gated: false), read 29 September 2026. https://huggingface.co/api/models/meta-models/Muse-Glimmer-30B At full precision the model targets 64 GB of VRAM. Meta compresses the weights to about 4 bits, which brings the language model under 20 GB.11
K-Quant-Dynamic needs 32 GB and gives up 0.2% on average across 15 benchmarks. The 17 GB version fits in 24 GB, at a cost of 1.0%.11
The config allows 131,072 positions, and Meta's blog gives 128K tokens as the default window.2244 Knowledge stops at 4 January 2026.11 Reasoning strength is set in the system prompt, from low up to xhigh.11
The Muse Spark models it was distilled from have no public weights.77Meta, Meta Model API documentation, which lists every Muse Spark version as an API model and Muse Glimmer as the self-hosted one, read 29 September 2026. https://dev.meta.ai/docs/models
Notes
-
Meta Superintelligence Labs, Muse Glimmer model card, Hugging Face, August 2026. https://huggingface.co/meta-models/Muse-Glimmer-30B 2 3 4 5 6 7 8 9 10 11 12 13 14
-
Meta, config.json in the official repository: num_hidden_layers 52, layer_types, sliding_window 2048, layer_rope_theta 0 on every full_attention layer, max_position_embeddings 131072. https://huggingface.co/meta-models/Muse-Glimmer-30B/blob/main/config.json. Arithmetic: 52 x 3/4 = 39 sliding layers and 52 x 1/4 = 13 full layers, as layer_types lists them. 2 3
-
Hugging Face, launch post for Muse Glimmer, 10 August 2026, which says the model was "released today". This is the earliest official date; Meta's developer blog post is dated 12 August 2026. https://huggingface.co/blog/muse-glimmer
-
Meta for Developers, "Build with Muse Glimmer", 12 August 2026: "the most permissive license we've used for an open model"; "a default context window of 128K tokens". https://dev.meta.ai/resources/blog/build-with-muse-glimmer 2
-
Meta, Llama 4 Community License Agreement, effective 5 April 2025. https://github.com/meta-llama/llama-models/blob/main/models/llama4/LICENSE
-
Hugging Face, repository metadata for meta-models/Muse-Glimmer-30B (gated: false), read 29 September 2026. https://huggingface.co/api/models/meta-models/Muse-Glimmer-30B
-
Meta, Meta Model API documentation, which lists every Muse Spark version as an API model and Muse Glimmer as the self-hosted one, read 29 September 2026. https://dev.meta.ai/docs/models
More from Meta
- Muse Spark 1.3TotalUndisclosedActiveUndisclosed
- Llama 4 MaverickTotal400BActive17B
- Llama 4 ScoutTotal109BActive17B