Skip the cover

Qwen

Qwen3.8-Flash

A 125B mixture of experts with 6B active per token, published as Qwen3.8-Flash-Next and built as a preview of the Qwen4 architecture.

Total
125B
Active
6B
Experts
512 routed + 1 shared10 per token
Layers
48
Attention
Gated DeltaNet and Qwen Sparse Attention
Context
1M tokens
Max output
131k tokens
Input
text, image, video
Output
text
Licence
Qwen Community License 1.0

Contents1 to 3

How it is built

Two names cover one model. The open weights are Qwen3.8-Flash-Next, and Qwen3.8-Flash is the API version built on them.11Qwen, Qwen3.8-Flash-Next model card, August 2026. SWE-bench Pro with the Claude Code harness, temperature 1.0, top_p 0.95 and a 256K context window. https://huggingface.co/Qwen/Qwen3.8-Flash-Next

N-gram embedding
A lookup table indexed by short runs of two or three tokens, read at layer 2.11

The main model has 125B parameters, with 6B active for each token.11 Qwen counts two further parts apart: an n-gram embedding of 51B parameters and a 4B module for multi-token prediction.11

N-gram embedding
A lookup table indexed by short runs of two or three tokens, read at layer 2.11

In every layer a token meets 10 routed experts out of 512, plus 1 shared expert.11 The 48 layers repeat a block of three Gated DeltaNet layers and one layer of Qwen Sparse Attention (QSA), 12 times over.11 Hidden states are 2,560 wide, and QSA layers use 24 query heads with 2 key-value heads.11 A vision encoder brings images and video into the same model.11

Fig. 1
Token

48 layers. In each one, a token goes to 10 of 512 routed experts and to 1 shared expert. The experts it meets here are simulated, because the real router depends on the trained weights.

What is new

Qwen presents the design as an experimental preview of the architecture planned for Qwen4.11 The weights and the blog post came out on 26 August 2026.22Qwen, Qwen3.8-Flash-Next blog post, 26 August 2026. Training "takes only about 1/9 as much" as Qwen3.7-Plus. https://qwen.ai/blog?id=qwen3.8-flash-next

QSA replaces the Gated Attention of earlier Qwen models. A small indexer picks context in micro-blocks rather than single tokens, with a budget of 512 blocks, or 2,048 tokens.11 Gated Residual widens the residual stream into 4 branches. An element-wise gate that depends on the data controls what is read, and each branch has a scalar gate for what is written.11

By Qwen's account, training cost about a ninth of what Qwen3.7-Plus took.22 On SWE-bench Pro the card gives it 62.5. The dense Qwen3.8-27B scores 61.7 in the same table, less than a point behind, and Qwen3.7-Plus 55.8.11 All three ran in Claude Code with a 256K context.11

Running it

The Qwen Community License 1.0 has no revenue threshold, unlike the licence of Qwen3.8-Max. Any company that runs a model-as-a-service or AI work assistant business needs a separate licence for commercial use, whatever its size.33Qwen, Qwen Community License 1.0, condition 2. https://huggingface.co/Qwen/Qwen3.8-Flash-Next/blob/main/LICENSE

Qwen says the n-gram table is easier to offload than experts, which suits accelerators with little memory.11 Past its native 262,144 tokens, the context stretches to 1,000,000 with YaRN.11 A request can turn thinking off.11

On QwenCloud the model runs as qwen3.8-flash, with 1M tokens of context and up to 131K of output.44QwenCloud, Qwen3.8-Flash model page, prices per million tokens as listed on 29 September 2026. https://www.qwencloud.com/models/qwen3.8-flash It costs $0.15 per million input tokens and $0.47 per million output tokens. Cached input costs $0.016.44

Notes

  1. Qwen, Qwen3.8-Flash-Next model card, August 2026. SWE-bench Pro with the Claude Code harness, temperature 1.0, top_p 0.95 and a 256K context window. https://huggingface.co/Qwen/Qwen3.8-Flash-Next 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16

  2. Qwen, Qwen3.8-Flash-Next blog post, 26 August 2026. Training "takes only about 1/9 as much" as Qwen3.7-Plus. https://qwen.ai/blog?id=qwen3.8-flash-next 2

  3. Qwen, Qwen Community License 1.0, condition 2. https://huggingface.co/Qwen/Qwen3.8-Flash-Next/blob/main/LICENSE

  4. QwenCloud, Qwen3.8-Flash model page, prices per million tokens as listed on 29 September 2026. https://www.qwencloud.com/models/qwen3.8-flash 2

More from Qwen

All models by Qwen