Qwen3.8-Omni-Flash
An API model from Qwen that reads audio and video as well as text and images, within a context of 1M tokens.
- Totale
- Non dichiarato
- Attivi
- Non dichiarato
- Esperti
- Non dichiarato
- Contesto
- 1M token
- Uscita massima
- 131k token
- Ingresso
- text, image, audio, video
- Uscita
- text
- Licenza
- Proprietary
- Pesi
- Solo API
Non dichiarato da Qwen
Indiceda 1 a 3
What is disclosed
Qwen3.8-Omni-Flash exists only on the API, as qwen3.8-omni-flash, and its weights are not published.11QwenCloud, Qwen3.8-Omni-Flash model page, prices per million tokens as listed on 29 September 2026. https://www.qwencloud.com/models/qwen3.8-omni-flash QwenCloud gives no parameter count and no expert count.11
According to the model page, it is built on the Qwen3.8-Flash-Next architecture.11 That open model, published as Qwen3.8-Flash, has 125B parameters with 6B active per token. Its layers alternate three Gated DeltaNet layers with one layer of Qwen Sparse Attention.22Qwen, Qwen3.8-Flash-Next model card, August 2026: 125B parameters with 6B active, hidden layout 12 x (3 x Gated DeltaNet, 1 x Qwen Sparse Attention). https://huggingface.co/Qwen/Qwen3.8-Flash-Next Qwen does not say whether Omni-Flash keeps those sizes.
Qwen calls the model natively omni-modal. Audio and video go in alongside text and images.11 It takes up to an hour of audio and video as input, a limit the announcement ties to meetings.33Alibaba Cloud, "Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery.", 20 September 2026. Up to one hour of audio-visual input; price comparisons are per hour of input, against Qwen3.5-Omni-Plus. https://www.alibabacloud.com/blog/603580 Speech recognition covers 74 languages and 39 Chinese dialects, Cantonese among them.33 The model also reads two-channel and four-channel spatial audio.11
The context window is 1M tokens. Input is capped at 991K, or 983K with thinking on, and output at 131K.11
What is new
Qwen announced the model on 18 September 2026.44Qwen, Qwen3.8-Omni-Flash blog post, 18 September 2026. https://qwen.ai/blog?id=qwen3.8-omni-flash Across 29 evaluations it reports an average more than 25% above Qwen3.5-Omni-Plus, its omni model of an earlier generation.33
Against Qwen3.5-Omni-Plus, an hour of audio input costs more than 98% less. An hour of audio with video costs more than 93% less.33
QwenCloud aims the model at agent work. Coding and GUI interaction are on its list, and so are video editing and dialogue over audio and video.11
Using it
A million input tokens cost $0.15 and a million output tokens $0.47. Cached input costs $0.016 per million.11 Qwen3.8-Flash, built on the same architecture, costs the same.55QwenCloud, Qwen3.8-Flash model page, as listed on 29 September 2026: $0.15 input and $0.47 output per million tokens. https://www.qwencloud.com/models/qwen3.8-flash
With thinking on, the same model id takes 8K fewer input tokens and keeps the 131K output limit.66QwenCloud, Qwen3.8-Omni-Flash model page. Arithmetic: 991K minus 983K = 8K. https://www.qwencloud.com/models/qwen3.8-omni-flash
The standard API returns text only.11 Yet the announcement describes speech generation in 29 languages and 7 dialects, Southern Min among them.33 Live audio and video streams go to a separate realtime version, Qwen3.8-Omni-Flash-Realtime, over WebSocket or WebRTC.33
Note
-
QwenCloud, Qwen3.8-Omni-Flash model page, prices per million tokens as listed on 29 September 2026. https://www.qwencloud.com/models/qwen3.8-omni-flash 2 3 4 5 6 7 8 9
-
Qwen, Qwen3.8-Flash-Next model card, August 2026: 125B parameters with 6B active, hidden layout 12 x (3 x Gated DeltaNet, 1 x Qwen Sparse Attention). https://huggingface.co/Qwen/Qwen3.8-Flash-Next
-
Alibaba Cloud, "Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery.", 20 September 2026. Up to one hour of audio-visual input; price comparisons are per hour of input, against Qwen3.5-Omni-Plus. https://www.alibabacloud.com/blog/603580 2 3 4 5 6
-
Qwen, Qwen3.8-Omni-Flash blog post, 18 September 2026. https://qwen.ai/blog?id=qwen3.8-omni-flash
-
QwenCloud, Qwen3.8-Flash model page, as listed on 29 September 2026: $0.15 input and $0.47 output per million tokens. https://www.qwencloud.com/models/qwen3.8-flash
-
QwenCloud, Qwen3.8-Omni-Flash model page. Arithmetic: 991K minus 983K = 8K. https://www.qwencloud.com/models/qwen3.8-omni-flash
Altri modelli di Qwen
- Qwen3.8-MaxTotale2,4TAttivi95B
- Qwen3.8-FlashTotale125BAttivi6B
- Qwen3.8-27BTotale27BAttivi27B