Qwen3.8-Max
Qwen's flagship, a 2.4T mixture of experts with 95B active per token, and the first Qwen-Max model released with open weights.
- Total
- 2.4T
- Active
- 95B
- Experts
- 512 routed + 1 shared10 per token
- Layers
- 92
- Attention
- Gated DeltaNet and Gated Attention
- Context
- 1M tokens
- Max output
- 131k tokens
- Input
- text, image, video
- Output
- text
- Licence
- Qwen3.8-Max License
- Weights
- Hugging Face
Contents1 to 3
How it is built
The open weights carry the name Qwen3.8-2.4T-A95B, which states both sizes: 2.4T parameters in total, 95B active for each token.11Qwen, Qwen3.8-2.4T-A95B model card, August 2026. Terminal Bench 2.1 with Claude Code (avg@10), a 5-hour timeout and max_tokens=131,072. SWE-bench Pro with the Claude Code harness, temperature 1.0, top_p 0.95 and a 256K context. reasoning_effort levels xhigh, medium and low. https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
Its 92 layers come in 23 groups of four, three Gated DeltaNet layers followed by one Gated Attention layer.11 Every layer ends in a mixture of 512 experts, and a token is routed to 10 of them plus 1 shared expert.11
The DeltaNet layers have 128 value heads and 16 query-key heads. In Gated Attention, 64 query heads share 4 key-value heads.11 The hidden size is 8,192, and training included multi-token prediction.11
What is new
Qwen had never released the weights of a Max model before.11 The API version was announced on 3 August 2026. The weights followed on 12 August.22Qwen, Qwen3.8-Max blog post, 3 August 2026. https://qwen.ai/blog?id=qwen3.8 33Qwen, Qwen3.8 repository README, news entry of 12 August 2026. https://github.com/QwenLM/Qwen3.8 The architecture builds on Qwen3.5.11
On Terminal Bench 2.1 Qwen reports 86.6, against 74.5 for Qwen3.7-Max, averaged over 10 attempts in the Claude Code harness.11 On SWE-bench Pro it reports 67.7, where Qwen3.7-Max had 60.6.11
Running it
The released weights read text only. They always think before they answer, and thinking cannot be turned off.11 The API version adds image and video input and a non-thinking mode, with a 1M context by default.1144QwenCloud, Qwen3.8-Max model page, prices per million tokens as listed on 29 September 2026. Input text, image and video; context 1M tokens. https://www.qwencloud.com/models/qwen3.8-max
reasoning_effort sets how deep the model thinks. The default is xhigh, the highest of three levels.11 For agentic work the card suggests 262,144 tokens for reasoning and 131,072 for the final answer.11
An FP8 checkpoint is published too. At one byte per parameter its weights come to about 2.4 TB.55Qwen, Qwen3.8-2.4T-A95B-FP8 model page. Arithmetic: 2.4T x 1 byte = 2.4 TB, before the KV cache and any tensors kept at higher precision. https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B-FP8
The licence allows commercial use on two conditions.66Qwen, Qwen3.8-Max License, conditions 1 and 2. Internal use that exposes nothing to third parties is exempt from condition 2. https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/blob/main/LICENSE A product with more than 100 million monthly active users, or US$20 million of monthly revenue, must display the model name. A company selling models as a service or AI work assistants needs a separate licence once its revenue passes US$50 million in twelve months.66
On QwenCloud, qwen3.8-max accepts up to 991K input tokens and returns up to 131K.44 Input costs $2 per million tokens, or $0.25 from the implicit cache. Output costs $6.44
Notes
-
Qwen, Qwen3.8-2.4T-A95B model card, August 2026. Terminal Bench 2.1 with Claude Code (avg@10), a 5-hour timeout and max_tokens=131,072. SWE-bench Pro with the Claude Code harness, temperature 1.0, top_p 0.95 and a 256K context. reasoning_effort levels xhigh, medium and low. https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B 2 3 4 5 6 7 8 9 10 11 12 13
-
Qwen, Qwen3.8-Max blog post, 3 August 2026. https://qwen.ai/blog?id=qwen3.8
-
Qwen, Qwen3.8 repository README, news entry of 12 August 2026. https://github.com/QwenLM/Qwen3.8
-
QwenCloud, Qwen3.8-Max model page, prices per million tokens as listed on 29 September 2026. Input text, image and video; context 1M tokens. https://www.qwencloud.com/models/qwen3.8-max 2 3
-
Qwen, Qwen3.8-2.4T-A95B-FP8 model page. Arithmetic: 2.4T x 1 byte = 2.4 TB, before the KV cache and any tensors kept at higher precision. https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B-FP8
-
Qwen, Qwen3.8-Max License, conditions 1 and 2. Internal use that exposes nothing to third parties is exempt from condition 2. https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/blob/main/LICENSE 2
More from Qwen
- Qwen3.8-Omni-FlashTotalUndisclosedActiveUndisclosed
- Qwen3.8-FlashTotal125BActive6B
- Qwen3.8-27BTotal27BActive27B