Skip the cover

Meta

Muse Glimmer 30B

Meta's open dense model of about 30B parameters, distilled from Muse Spark to run agents on a Mac or a single consumer GPU.

Total
29.6B
Active
29.6B
Experts
Nonedense model
Layers
52
Attention
Local and global 3:1, gated GQA
Context
131k tokens
Input
text, image
Output
text
Licence
Apache 2.0

Contents1 to 3

How it is built

Muse Glimmer counts about 29.6B parameters, and that figure includes the vision encoder.11Meta Superintelligence Labs, Muse Glimmer model card, Hugging Face, August 2026. https://huggingface.co/meta-models/Muse-Glimmer-30B The encoder alone is a ViT-G/14 of about 1.8B. Meta's card calls the model "distilled from Muse Spark" and built "for autonomous agentic tasks on consumer hardware".11

Sliding window
A local layer lets each token attend only to a fixed span of recent tokens. A global layer sees the whole context.

It is a dense transformer. The language model has 52 layers with a hidden size of 6,656.11 Three local layers come before every global one, which gives 39 local and 13 global (52 x 3/4 and 52 x 1/4).22Meta, config.json in the official repository: num_hidden_layers 52, layer_types, sliding_window 2048, layer_rope_theta 0 on every full_attention layer, max_position_embeddings 131072. https://huggingface.co/meta-models/Muse-Glimmer-30B/blob/main/config.json. Arithmetic: 52 x 3/4 = 39 sliding layers and 52 x 1/4 = 13 full layers, as layer_types lists them. The local layers attend over a sliding window of 2,048 tokens.11

Sliding window
A local layer lets each token attend only to a fixed span of recent tokens. A global layer sees the whole context.

RoPE, with a base of 500,000, runs only in the local layers. In the config the global layers get no rotary position at all.22 Each layer has 32 query heads sharing 2 key-value heads, with gated attention.11 The feed-forward block is SwiGLU, 19,968 wide.11

Fig. 1

52 layers, all dense.

What is new

The weights went public on Hugging Face on 10 August 2026, under Apache 2.0.33Hugging Face, launch post for Muse Glimmer, 10 August 2026, which says the model was "released today". This is the earliest official date; Meta's developer blog post is dated 12 August 2026. https://huggingface.co/blog/muse-glimmer Meta's developer blog calls it the most permissive licence Meta has used for an open model.44Meta for Developers, "Build with Muse Glimmer", 12 August 2026: "the most permissive license we've used for an open model"; "a default context window of 128K tokens". https://dev.meta.ai/resources/blog/build-with-muse-glimmer Llama 4 had come under the Llama 4 Community License.55Meta, Llama 4 Community License Agreement, effective 5 April 2025. https://github.com/meta-llama/llama-models/blob/main/models/llama4/LICENSE

Speculative decoding
A small model drafts tokens ahead, and the large model checks them. It keeps the tokens it would have written itself and corrects the rest.

For speculative decoding the model ships with a small DFlash drafter. It proposes 16 tokens in one forward pass, and the main model verifies them in parallel.11

Speculative decoding
A small model drafts tokens ahead, and the large model checks them. It keeps the tokens it would have written itself and corrects the rest.

Meta measured the drafter with its 17 GB quantization at batch size 1. On an NVIDIA RTX 5090 under llama.cpp, generation went from 74.9 to 233.4 tokens per second.11 An Apple M5 Max under ExecuTorch went from 26.6 to 50.2.11

The comparison on the card is with Gemma 4 31B and Qwen3.6-27B, both thinking. At high reasoning Muse Glimmer scores 75.5 on MCP Atlas (Public), where the other two score 54.2 and 62.5. On SWE-Bench Verified it trails Qwen, 76.0 against 77.2, with Gemma at 66.6.11

Running it

The repository has no access gate.66Hugging Face, repository metadata for meta-models/Muse-Glimmer-30B (gated: false), read 29 September 2026. https://huggingface.co/api/models/meta-models/Muse-Glimmer-30B At full precision the model targets 64 GB of VRAM. Meta compresses the weights to about 4 bits, which brings the language model under 20 GB.11

K-Quant-Dynamic needs 32 GB and gives up 0.2% on average across 15 benchmarks. The 17 GB version fits in 24 GB, at a cost of 1.0%.11

The config allows 131,072 positions, and Meta's blog gives 128K tokens as the default window.2244 Knowledge stops at 4 January 2026.11 Reasoning strength is set in the system prompt, from low up to xhigh.11

The Muse Spark models it was distilled from have no public weights.77Meta, Meta Model API documentation, which lists every Muse Spark version as an API model and Muse Glimmer as the self-hosted one, read 29 September 2026. https://dev.meta.ai/docs/models

Notes

  1. Meta Superintelligence Labs, Muse Glimmer model card, Hugging Face, August 2026. https://huggingface.co/meta-models/Muse-Glimmer-30B 2 3 4 5 6 7 8 9 10 11 12 13 14

  2. Meta, config.json in the official repository: num_hidden_layers 52, layer_types, sliding_window 2048, layer_rope_theta 0 on every full_attention layer, max_position_embeddings 131072. https://huggingface.co/meta-models/Muse-Glimmer-30B/blob/main/config.json. Arithmetic: 52 x 3/4 = 39 sliding layers and 52 x 1/4 = 13 full layers, as layer_types lists them. 2 3

  3. Hugging Face, launch post for Muse Glimmer, 10 August 2026, which says the model was "released today". This is the earliest official date; Meta's developer blog post is dated 12 August 2026. https://huggingface.co/blog/muse-glimmer

  4. Meta for Developers, "Build with Muse Glimmer", 12 August 2026: "the most permissive license we've used for an open model"; "a default context window of 128K tokens". https://dev.meta.ai/resources/blog/build-with-muse-glimmer 2

  5. Meta, Llama 4 Community License Agreement, effective 5 April 2025. https://github.com/meta-llama/llama-models/blob/main/models/llama4/LICENSE

  6. Hugging Face, repository metadata for meta-models/Muse-Glimmer-30B (gated: false), read 29 September 2026. https://huggingface.co/api/models/meta-models/Muse-Glimmer-30B

  7. Meta, Meta Model API documentation, which lists every Muse Spark version as an API model and Muse Glimmer as the self-hosted one, read 29 September 2026. https://dev.meta.ai/docs/models

More from Meta

All models by Meta