{"slug": "qwen-3-8-flash-next-config-llama-cpp", "title": "Qwen 3.8 Flash Next config LLama.cpp", "summary": "A developer benchmarked the Qwen3.8-Flash-Next GGUF quant (AD-4.27bpw Q4_K_M, 94.5 GB across 33 shards) on a single RTX 5060 Ti 16GB with 64GB DDR4, publishing llama.cpp server configurations that sustain 20–32 t/s decode at up to 131k context and roughly 16 t/s at 262k. The writeup derives an empirical rule that ubatch × context must stay under about 540 million, above which decode collapses (8.9 t/s at 671M) and generation fails outright at ~940M, and notes that ngram-mod speculative decoding makes throughput draft-acceptance-bound, ranging from 12.1 to 36.3 t/s on identical configs.", "body_md": "**Validated on:** RTX 5060 Ti 16GB · i5-14500 · 64GB DDR4-3200 · NVMe Gen4 · llama.cpp `0.5.0-dev` (needs `qwen4exp` arch + ngram spec support, any build after ~Sep 27 2026). All numbers from a two-round A/B-A/B bench, 2026-09-28.\n\n```\nhuggingface.co/AtomicChat/Qwen3.8-Flash-Next-GGUF\n  → Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64/   # 94.5 GB, 33 shards; shard 2 = 38 GB n-gram table (keep on NVMe!)\n  → mmproj-Qwen3.8-Flash-Next-F16.gguf           # 0.9 GB vision encoder\n```\n\n**Daily driver — 120k context, vision included:**\n\n```\nllama-server \\\n  -m Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64/Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64-00001-of-00033.gguf \\\n  --mmproj mmproj-Qwen3.8-Flash-Next-F16.gguf --no-mmproj-offload \\\n  --host 127.0.0.1 --port 8014 \\\n  -c 122880 \\\n  --fit on --fit-target 128 \\\n  -ctk q8_0 -ctv q8_0 \\\n  -fa on \\\n  --jinja \\\n  -np 1 -b 4096 -ub 4096 \\\n  -t 20 --threads-batch 20 --prio 2 \\\n  --lazy-mode on \\\n  --spec-type ngram-mod \\\n  --spec-ngram-mod-n-match 60 --spec-ngram-mod-n-min 12 --spec-ngram-mod-n-max 24 \\\n  --cache-reuse 256 \\\n  --reasoning-format deepseek\n```\n\n**Fast preset:** identical command with `-c 81920`. **Full window (262k):** `-c 262144` **and** switch to `-b 2048 -ub 1024 -t 10 --threads-batch 12` (why below).\n\n| Preset | 16k | 40k | 62k | 78k | 114k | Prefill | \n|---|---|---|---|---|---|---|\n| 80k (fast) | 30.8 | 29.6 | 27.7 | 27.3 | — | 310–360 t/s | \n| 120k (daily) | 26.2 | 28.9 | — | 12–36* | 22.5 | 300–360 t/s | \n| 260k (full) | — | — | — | — | ~16 @222k+ | ~190–200 t/s | \n\n* decode is **draft-acceptance-bound**: `ngram-mod` drafts from the model's own n-gram table, and acceptance depends on content. I measured the *identical* rung twice at 12.1 t/s (12.11/12.06) and a different-content run at 36.3 t/s — same config, same depth. Repetitive text/code drafts well (expect the high end); hostile content can halve it. Quote bands, not points.\n\nCompute buffers scale with `ubatch`, KV with `ctx`, and together they squeeze experts off the 16GB GPU. Every measured point fits one empirical law:\n\n**Keep `ub × ctx` under ~540 million.** At ~671M decode collapses; at ~940M generation breaks outright.\n\n| Your target ctx | Use | Product | Measured | \n|---|---|---|---|\n| ≤ 131k | `-b 4096 -ub 4096 -t 20 -tb 20` | ≤537M | 20–32 t/s ✓ | \n| 131k–180k | `-b 2048 -ub 2048 -t 20 -tb 20` | ~410M @163k | (untested corner — should hold by the law) | \n| up to 262k | `-b 2048 -ub 1024 -t 10 -tb 12` | 268M | ~16 t/s ✓ | \n| never | `-ub 8192` | 671M even @82k | 8.9 t/s ✗ | \n\nSpecifically: `-ub 4096` at 163k context decodes at 10.5 t/s and at 229k **fails to generate at all** — if you need the full window, the small ubatch is mandatory, and you trade ~45% of prefill speed (~200 vs ~340 t/s) for the range.\n\n- `--lazy-mode on` + mmap (default): the 38GB n-gram shard stays on SSD; rows fault in on demand. This is what makes 64GB RAM work. Requires mmap — never combine with`--load-mode none` .\n- `--no-mmproj-offload` : vision encoder on CPU =**free** ; on GPU it costs −19% decode for faster image encoding. A few extra seconds per image.\n- `-t 20` on all cores including E-cores: saturates ~7–10 cores of work.**Never pin to P-cores only** — strict P-core affinity halved decode (6 physical cores + lost E-core memory streams).\n- `--fit on --fit-target 128` : auto GPU/CPU layer split, beats manual`-ncmoe` by ~3%.\n- CPU governor: irrelevant (memory-bandwidth-bound) — leave powersave.\n\n- First request after load runs ~30s slow (NVMe expert page-in; 15 t/s prefill cold → 300+ warm).\n- GPU util oscillates 0%↔busy ~1/sec during decode — GPU waiting on CPU-resident experts. Normal.\n- ~46–50GB of resident mmap pages; under RAM pressure the kernel evicts/re-faults (dips, not crashes). In Docker, don't cap the container below ~52GB.\n- VRAM steady ~15.7GB; q8_0 KV; no MTP draft head needed — the n-gram table is the drafter.\n\n**Avoid:** `-ot 'exps=CPU'` (~4 t/s), `-ngl 99` alone (OOM), `--load-mode none` (breaks lazy-mode; needs ~93GB), EXL3/vLLM (won't fit), `-ub 8192` (see table).", "url": "https://wpnews.pro/news/qwen-3-8-flash-next-config-llama-cpp", "canonical_source": "https://gist.github.com/sdurnov/3c8586c5017c6ac3fa9641e8653754f8", "published_at": "2026-09-28 15:04:57+00:00", "updated_at": "2026-09-29 08:18:31.880988+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "mlops", "ai-tools", "developer-tools"], "entities": ["Qwen", "llama.cpp", "RTX 5060 Ti", "AtomicChat", "Hugging Face", "GGUF"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/qwen-3-8-flash-next-config-llama-cpp", "markdown": "https://wpnews.pro/news/qwen-3-8-flash-next-config-llama-cpp.md", "text": "https://wpnews.pro/news/qwen-3-8-flash-next-config-llama-cpp.txt", "jsonld": "https://wpnews.pro/news/qwen-3-8-flash-next-config-llama-cpp.jsonld"}}