{"slug": "alibabas-qwen-team-releases-qwen3-8-flash-next-a-125b-multimodal-moe-with-6b-the", "title": "Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture", "summary": "Alibaba's Qwen team released Qwen3.8-Flash-Next, an open-weight multimodal Mixture-of-Experts model with 125B backbone parameters and 6B active per token, previewing the Qwen4 architecture. The model includes a 51B N-gram embedding table and a 4B multi-token prediction module, and reports a 1/9 training cost compared to Qwen3.7-Plus. The FP8 checkpoint is 172.78 GiB, and the model introduces four architectural changes: Gated DeltaNet and Qwen Sparse Attention hybrid, Gated Residual, N-gram Embedding, and the Muon optimizer.", "body_md": "We look at Qwen3.8-Flash-Next, Alibaba's open-weight multimodal Mixture-of-Experts model and an early preview of the Qwen4 architecture. We break down where the 180B parameters actually sit: a 125B backbone, a 51B N-gram embedding table, and a 4B multi-token prediction module, with only 6B active per token. We walk through the four architectural changes — the Gated DeltaNet and Qwen Sparse Attention hybrid, Gated Residual, N-gram Embedding, and the Muon optimizer. We also cover the benchmark results, the reported 1/9 training cost against Qwen3.7-Plus, and what self-hosting a 172.78 GiB FP8 checkpoint really demands.\n\nThe post [Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture](https://www.marktechpost.com/2026/08/26/alibabas-qwen-team-releases-qwen3-8-flash-next-a-125b-multimodal-moe-with-6b-active-parameters-previewing-the-qwen4-architecture/) appeared first on [MarkTechPost](https://www.marktechpost.com).", "url": "https://wpnews.pro/news/alibabas-qwen-team-releases-qwen3-8-flash-next-a-125b-multimodal-moe-with-6b-the", "canonical_source": "https://www.marktechpost.com/2026/08/26/alibabas-qwen-team-releases-qwen3-8-flash-next-a-125b-multimodal-moe-with-6b-active-parameters-previewing-the-qwen4-architecture/", "published_at": "2026-08-26 15:20:00+00:00", "updated_at": "2026-08-26 15:45:27.202780+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "generative-ai", "ai-research"], "entities": ["Alibaba", "Qwen", "Qwen3.8-Flash-Next", "Qwen4", "Qwen3.7-Plus", "MarkTechPost"], "alternates": {"html": "https://wpnews.pro/news/alibabas-qwen-team-releases-qwen3-8-flash-next-a-125b-multimodal-moe-with-6b-the", "markdown": "https://wpnews.pro/news/alibabas-qwen-team-releases-qwen3-8-flash-next-a-125b-multimodal-moe-with-6b-the.md", "text": "https://wpnews.pro/news/alibabas-qwen-team-releases-qwen3-8-flash-next-a-125b-multimodal-moe-with-6b-the.txt", "jsonld": "https://wpnews.pro/news/alibabas-qwen-team-releases-qwen3-8-flash-next-a-125b-multimodal-moe-with-6b-the.jsonld"}}