Skip the cover

Qwen

Qwen3.8-27B

A dense vision-language model of 27B parameters, and the only open Qwen3.8 release under Apache 2.0.

Total
27B
Active
27B
Experts
Nonedense model
Layers
64
Attention
Gated DeltaNet and Gated Attention
Context
262k tokens
Input
text, image, video
Output
text
Licence
Apache 2.0

Contents1 to 3

How it is built

Qwen3.8-27B is the dense model of the family. All 27B parameters work on every token.11Qwen, Qwen3.8-27B model card, August 2026. SWE-bench Pro and DeepSWE 1.1 with the Claude Code harness, temperature 1.0, top_p 0.95 and a 256K context window; the Terminal Bench 2.1 row is labelled "Terminus"; video understanding "from STEM diagrams and documents to hour-scale videos". https://huggingface.co/Qwen/Qwen3.8-27B Qwen3.8-Max and Qwen3.8-Flash, the two larger open models, send each token through a handful of experts instead.22Qwen, Qwen3.8-2.4T-A95B model card: 2.4T parameters with 95B active, layout 23 x (3 x Gated DeltaNet, 1 x Gated Attention), thinking that cannot be disabled, Qwen3.8-Max License. https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B 33Qwen, Qwen3.8-Flash-Next model card: 125B parameters with 6B active, Qwen Community License 1.0. https://huggingface.co/Qwen/Qwen3.8-Flash-Next

Dense model
A model that uses all of its weights for every token, with no routing to experts.

The layer pattern is theirs. The 64 layers form 16 blocks of three Gated DeltaNet layers and one Gated Attention layer.112233 The hidden size is 5,120, and each feed-forward block is 17,408 wide. Attention uses 24 query heads and 4 key-value heads.11 The padded vocabulary holds 248,320 tokens, and training added multi-token prediction.11

Dense model
A model that uses all of its weights for every token, with no routing to experts.

A vision encoder of 27 layers handles images and video.44Qwen, Qwen3.8-27B config.json (text dtype bfloat16, vision encoder depth 27). Arithmetic: 27B x 2 bytes = 54 GB, before the vision encoder and the KV cache. https://huggingface.co/Qwen/Qwen3.8-27B/blob/main/config.json Qwen calls the result a native vision-language model and says it follows videos an hour long.11

Fig. 1

64 layers, all dense.

What is new

Qwen released it on 14 August 2026, two days after the weights of the 2.4T model.55Qwen, Qwen3.8 repository README, news entries of 12 August 2026 (Qwen3.8-2.4T-A95B weights) and 14 August 2026 (Qwen3.8-27B). https://github.com/QwenLM/Qwen3.8

Against Qwen3.6-27B, the card reports 61.7 on SWE-bench Pro where the older model had 53.5, both in the Claude Code harness at temperature 1.0 with a 256K context.11 With the Terminus harness, Terminal Bench 2.1 went from 63.4 to 73.0.11 NL2Repo-Bench asks for a whole repository from a written description. There the score is 42.3, against 36.2.11 On DeepSWE 1.1 the older model resolved 13.3 and the new one 42.2.11

Thinking is on by default and can be turned off for a single request.11 The open 2.4T model has no such switch.22 Depth is tuned with reasoning_effort, and preserve_thinking keeps the reasoning of earlier turns in context.11

Running it

This is the one open Qwen3.8 model under Apache 2.0.11 The two larger ones carry Qwen's own licences, with conditions on commercial use.2233

In BF16, at two bytes per parameter, the language model's weights take about 54 GB.44 The FP8 checkpoint halves that to about 27 GB.66Qwen, Qwen3.8-27B-FP8 model page. Arithmetic: 27B x 1 byte = 27 GB. https://huggingface.co/Qwen/Qwen3.8-27B-FP8

YaRN
A method that stretches the position encoding so a model can read past the context length it was trained on.

The context is 262,144 tokens natively and 1,000,000 with YaRN scaling.11 Qwen has announced a hosted version with a 1M context by default. The card still lists it as coming soon.11

YaRN
A method that stretches the position encoding so a model can read past the context length it was trained on.

Notes

  1. Qwen, Qwen3.8-27B model card, August 2026. SWE-bench Pro and DeepSWE 1.1 with the Claude Code harness, temperature 1.0, top_p 0.95 and a 256K context window; the Terminal Bench 2.1 row is labelled "Terminus"; video understanding "from STEM diagrams and documents to hour-scale videos". https://huggingface.co/Qwen/Qwen3.8-27B 2 3 4 5 6 7 8 9 10 11 12 13 14

  2. Qwen, Qwen3.8-2.4T-A95B model card: 2.4T parameters with 95B active, layout 23 x (3 x Gated DeltaNet, 1 x Gated Attention), thinking that cannot be disabled, Qwen3.8-Max License. https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B 2 3 4

  3. Qwen, Qwen3.8-Flash-Next model card: 125B parameters with 6B active, Qwen Community License 1.0. https://huggingface.co/Qwen/Qwen3.8-Flash-Next 2 3

  4. Qwen, Qwen3.8-27B config.json (text dtype bfloat16, vision encoder depth 27). Arithmetic: 27B x 2 bytes = 54 GB, before the vision encoder and the KV cache. https://huggingface.co/Qwen/Qwen3.8-27B/blob/main/config.json 2

  5. Qwen, Qwen3.8 repository README, news entries of 12 August 2026 (Qwen3.8-2.4T-A95B weights) and 14 August 2026 (Qwen3.8-27B). https://github.com/QwenLM/Qwen3.8

  6. Qwen, Qwen3.8-27B-FP8 model page. Arithmetic: 27B x 1 byte = 27 GB. https://huggingface.co/Qwen/Qwen3.8-27B-FP8

More from Qwen

All models by Qwen