# Qwen3.8-Flash-Next

> Source: <https://tokenstead.ai/models/qwen3-8-flash-next>
> Published: 2026-08-26 15:40:39+00:00

# Qwen3.8-Flash-Next

MoE enthusiast**125B total MoE, only 6B active per token** - the first open-weight preview of the Qwen4 architecture, released 2026-08-26. The active-parameter count is the headline: a fraction of a frontier model’s, yet it tops Qwen3.7-Plus-Base on 8 of 14 benchmarks. On top of the 6B active: **51B n-gram embeddings** (deterministic host-memory lookups, no per-token compute) and **4B MTP**. qwen-community-1.0 license on HuggingFace at `Qwen/Qwen3.8-Flash-Next`

.

-
**Hybrid attention (GDN + QSA):** Gated DeltaNet compresses history; Qwen Sparse Attention (QSA) does micro-block context selection via a lightweight indexer. At 1M context QSA’s kernel is up to 7.6x faster prefill / 4.9x faster decode vs prior attention; with 90% prefix-cache hits, 8.6x the prefill throughput of Qwen3.7-Plus. -
**Gated Residual:** 4-branch residual stream with data-dependent read gating + per-branch scalar write gating - better cross-layer flow, stronger training stability. -
**Optimization:** Muon optimizer (improved orthogonalization, Muon/AdamW split), no batch-size warmup, refitted scaling laws. -
**Multimodal:** text + image + video in, text out. 262K native context, extensible to ~1M via YaRN.

**Benchmarks (Qwen self-reported):** DeepSWE 58.7, SWE-bench Pro 62.5, SWE-bench Multilingual 81.0, GPQA-Diamond 91.7, LiveCodeBench v6 91.9, IFBench 81.3, Toolathlon Verified 73.5, HLE 35.9. Strongest open-source agentic-coded results at this active-param scale.

**Local-run status:** Unsloth’s Dynamic 3.0 GGUF landed the same day at `unsloth/Qwen3.8-Flash-Next-GGUF`

. **UD-IQ1_S** (72.5GB, 1-bit dynamic) is the first published quant - ~78GB to run in RAM. The 51B n-gram embeddings live in host memory (deterministic lookups), and the 6B-active MoE + GDN/QSA hybrid attention keeps the KV cache small, so the file is the footprint that matters. More quants are still uploading (repo marked WIP). Treat the 6B-active efficiency claims as vendor-reported until community replication. Note this is the experimental Flash-Next checkpoint; the production **Qwen3.8-Flash** service (1M default context, built-in tools) is a separate hosted product on Qwen Cloud.

- 125.0B
- 262k
- qwen community 1.0
- 🇨🇳 China
- Aug 2026

## Scores

## Run it locally

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

### Can you run it? - reference rigs

| Rig | UD-IQ1_S |
|---|---|
| NVIDIA Jetson Orin NX 16GB |
|

[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

## Download options

## Or run it in the cloud

No per-token API provider pricing tracked for Qwen3.8-Flash-Next yet.
For flagship list prices, see the
[calculator](/calculator).

[See who runs Alibaba in production →](/adoption/alibaba)

## Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.
