Qwen3.8-Flash-Next Alibaba released Qwen3.8-Flash-Next, the first open-weight preview of the Qwen4 architecture, on 2026-08-26, featuring 125B total parameters with only 6B active per token and topping Qwen3.7-Plus-Base on 8 of 14 benchmarks. The model introduces hybrid attention (GDN + QSA), a gated residual stream, and multimodal input support, with self-reported scores including DeepSWE 58.7, SWE-bench Pro 62.5, and GPQA-Diamond 91.7. Unsloth published a Dynamic 3.0 GGUF quant the same day, with the UD-IQ1_S (72.5GB) as the first quant, though efficiency claims remain vendor-reported until community replication. Qwen3.8-Flash-Next MoE enthusiast 125B total MoE, only 6B active per token - the first open-weight preview of the Qwen4 architecture, released 2026-08-26. The active-parameter count is the headline: a fraction of a frontier model’s, yet it tops Qwen3.7-Plus-Base on 8 of 14 benchmarks. On top of the 6B active: 51B n-gram embeddings deterministic host-memory lookups, no per-token compute and 4B MTP . qwen-community-1.0 license on HuggingFace at Qwen/Qwen3.8-Flash-Next . - Hybrid attention GDN + QSA : Gated DeltaNet compresses history; Qwen Sparse Attention QSA does micro-block context selection via a lightweight indexer. At 1M context QSA’s kernel is up to 7.6x faster prefill / 4.9x faster decode vs prior attention; with 90% prefix-cache hits, 8.6x the prefill throughput of Qwen3.7-Plus. - Gated Residual: 4-branch residual stream with data-dependent read gating + per-branch scalar write gating - better cross-layer flow, stronger training stability. - Optimization: Muon optimizer improved orthogonalization, Muon/AdamW split , no batch-size warmup, refitted scaling laws. - Multimodal: text + image + video in, text out. 262K native context, extensible to ~1M via YaRN. Benchmarks Qwen self-reported : DeepSWE 58.7, SWE-bench Pro 62.5, SWE-bench Multilingual 81.0, GPQA-Diamond 91.7, LiveCodeBench v6 91.9, IFBench 81.3, Toolathlon Verified 73.5, HLE 35.9. Strongest open-source agentic-coded results at this active-param scale. Local-run status: Unsloth’s Dynamic 3.0 GGUF landed the same day at unsloth/Qwen3.8-Flash-Next-GGUF . UD-IQ1 S 72.5GB, 1-bit dynamic is the first published quant - ~78GB to run in RAM. The 51B n-gram embeddings live in host memory deterministic lookups , and the 6B-active MoE + GDN/QSA hybrid attention keeps the KV cache small, so the file is the footprint that matters. More quants are still uploading repo marked WIP . Treat the 6B-active efficiency claims as vendor-reported until community replication. Note this is the experimental Flash-Next checkpoint; the production Qwen3.8-Flash service 1M default context, built-in tools is a separate hosted product on Qwen Cloud. - 125.0B - 262k - qwen community 1.0 - 🇨🇳 China - Aug 2026 Scores Run it locally Per-quant memory needs and a static "can you run it?" reference - no rig entry required Can you run it? - reference rigs | Rig | UD-IQ1 S | |---|---| | NVIDIA Jetson Orin NX 16GB | | no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast =20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark. Download options Or run it in the cloud No per-token API provider pricing tracked for Qwen3.8-Flash-Next yet. For flagship list prices, see the calculator /calculator . See who runs Alibaba in production → /adoption/alibaba Inference cost over time Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.