{"slug": "qwen3-8-flash-next", "title": "Qwen3.8-Flash-Next", "summary": "Alibaba released Qwen3.8-Flash-Next, the first open-weight preview of the Qwen4 architecture, on 2026-08-26, featuring 125B total parameters with only 6B active per token and topping Qwen3.7-Plus-Base on 8 of 14 benchmarks. The model introduces hybrid attention (GDN + QSA), a gated residual stream, and multimodal input support, with self-reported scores including DeepSWE 58.7, SWE-bench Pro 62.5, and GPQA-Diamond 91.7. Unsloth published a Dynamic 3.0 GGUF quant the same day, with the UD-IQ1_S (72.5GB) as the first quant, though efficiency claims remain vendor-reported until community replication.", "body_md": "# Qwen3.8-Flash-Next\n\nMoE enthusiast**125B total MoE, only 6B active per token** - the first open-weight preview of the Qwen4 architecture, released 2026-08-26. The active-parameter count is the headline: a fraction of a frontier model’s, yet it tops Qwen3.7-Plus-Base on 8 of 14 benchmarks. On top of the 6B active: **51B n-gram embeddings** (deterministic host-memory lookups, no per-token compute) and **4B MTP**. qwen-community-1.0 license on HuggingFace at `Qwen/Qwen3.8-Flash-Next`\n\n.\n\n-\n**Hybrid attention (GDN + QSA):** Gated DeltaNet compresses history; Qwen Sparse Attention (QSA) does micro-block context selection via a lightweight indexer. At 1M context QSA’s kernel is up to 7.6x faster prefill / 4.9x faster decode vs prior attention; with 90% prefix-cache hits, 8.6x the prefill throughput of Qwen3.7-Plus. -\n**Gated Residual:** 4-branch residual stream with data-dependent read gating + per-branch scalar write gating - better cross-layer flow, stronger training stability. -\n**Optimization:** Muon optimizer (improved orthogonalization, Muon/AdamW split), no batch-size warmup, refitted scaling laws. -\n**Multimodal:** text + image + video in, text out. 262K native context, extensible to ~1M via YaRN.\n\n**Benchmarks (Qwen self-reported):** DeepSWE 58.7, SWE-bench Pro 62.5, SWE-bench Multilingual 81.0, GPQA-Diamond 91.7, LiveCodeBench v6 91.9, IFBench 81.3, Toolathlon Verified 73.5, HLE 35.9. Strongest open-source agentic-coded results at this active-param scale.\n\n**Local-run status:** Unsloth’s Dynamic 3.0 GGUF landed the same day at `unsloth/Qwen3.8-Flash-Next-GGUF`\n\n. **UD-IQ1_S** (72.5GB, 1-bit dynamic) is the first published quant - ~78GB to run in RAM. The 51B n-gram embeddings live in host memory (deterministic lookups), and the 6B-active MoE + GDN/QSA hybrid attention keeps the KV cache small, so the file is the footprint that matters. More quants are still uploading (repo marked WIP). Treat the 6B-active efficiency claims as vendor-reported until community replication. Note this is the experimental Flash-Next checkpoint; the production **Qwen3.8-Flash** service (1M default context, built-in tools) is a separate hosted product on Qwen Cloud.\n\n- 125.0B\n- 262k\n- qwen community 1.0\n- 🇨🇳 China\n- Aug 2026\n\n## Scores\n\n## Run it locally\n\nPer-quant memory needs and a static \"can you run it?\" reference - no rig entry required\n\n### Can you run it? - reference rigs\n\n| Rig | UD-IQ1_S |\n|---|---|\n| NVIDIA Jetson Orin NX 16GB |\n|\n\n[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.\n\n## Download options\n\n## Or run it in the cloud\n\nNo per-token API provider pricing tracked for Qwen3.8-Flash-Next yet.\nFor flagship list prices, see the\n[calculator](/calculator).\n\n[See who runs Alibaba in production →](/adoption/alibaba)\n\n## Inference cost over time\n\nData accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.", "url": "https://wpnews.pro/news/qwen3-8-flash-next", "canonical_source": "https://tokenstead.ai/models/qwen3-8-flash-next", "published_at": "2026-08-26 15:40:39+00:00", "updated_at": "2026-08-26 15:44:41.596330+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-products", "ai-infrastructure"], "entities": ["Alibaba", "Qwen3.8-Flash-Next", "Qwen4", "Qwen3.7-Plus-Base", "Unsloth", "HuggingFace", "Qwen Cloud"], "alternates": {"html": "https://wpnews.pro/news/qwen3-8-flash-next", "markdown": "https://wpnews.pro/news/qwen3-8-flash-next.md", "text": "https://wpnews.pro/news/qwen3-8-flash-next.txt", "jsonld": "https://wpnews.pro/news/qwen3-8-flash-next.jsonld"}}