MoE enthusiast125B total MoE, only 6B active per token - the first open-weight preview of the Qwen4 architecture, released 2026-08-26. The active-parameter count is the headline: a fraction of a frontier model’s, yet it tops Qwen3.7-Plus-Base on 8 of 14 benchmarks. On top of the 6B active: 51B n-gram embeddings (deterministic host-memory lookups, no per-token compute) and 4B MTP. qwen-community-1.0 license on HuggingFace at Qwen/Qwen3.8-Flash-Next
.
Hybrid attention (GDN + QSA): Gated DeltaNet compresses history; Qwen Sparse Attention (QSA) does micro-block context selection via a lightweight indexer. At 1M context QSA’s kernel is up to 7.6x faster prefill / 4.9x faster decode vs prior attention; with 90% prefix-cache hits, 8.6x the prefill throughput of Qwen3.7-Plus. - Gated Residual: 4-branch residual stream with data-dependent read gating + per-branch scalar write gating - better cross-layer flow, stronger training stability. - Optimization: Muon optimizer (improved orthogonalization, Muon/AdamW split), no batch-size warmup, refitted scaling laws. - Multimodal: text + image + video in, text out. 262K native context, extensible to ~1M via YaRN.
Benchmarks (Qwen self-reported): DeepSWE 58.7, SWE-bench Pro 62.5, SWE-bench Multilingual 81.0, GPQA-Diamond 91.7, LiveCodeBench v6 91.9, IFBench 81.3, Toolathlon Verified 73.5, HLE 35.9. Strongest open-source agentic-coded results at this active-param scale.
Local-run status: Unsloth’s Dynamic 3.0 GGUF landed the same day at unsloth/Qwen3.8-Flash-Next-GGUF
. UD-IQ1_S (72.5GB, 1-bit dynamic) is the first published quant - ~78GB to run in RAM. The 51B n-gram embeddings live in host memory (deterministic lookups), and the 6B-active MoE + GDN/QSA hybrid attention keeps the KV cache small, so the file is the footprint that matters. More quants are still up (repo marked WIP). Treat the 6B-active efficiency claims as vendor-reported until community replication. Note this is the experimental Flash-Next checkpoint; the production Qwen3.8-Flash service (1M default context, built-in tools) is a separate hosted product on Qwen Cloud.
- 125.0B
- 262k
- qwen community 1.0
- 🇨🇳 China
- Aug 2026
Scores #
Run it locally #
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
Can you run it? - reference rigs
| Rig | UD-IQ1_S |
|---|---|
| NVIDIA Jetson Orin NX 16GB | |
no -> cloudno -> cloudno -> cloudFit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
Download options #
Or run it in the cloud #
No per-token API provider pricing tracked for Qwen3.8-Flash-Next yet.
For flagship list prices, see the
[calculator](/calculator).
[See who runs Alibaba in production →](/adoption/alibaba)
Inference cost over time #
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.