# Qwen 3.8 Flash Next config LLama.cpp

> Source: <https://gist.github.com/sdurnov/3c8586c5017c6ac3fa9641e8653754f8>
> Published: 2026-09-28 15:04:57+00:00

**Validated on:** RTX 5060 Ti 16GB · i5-14500 · 64GB DDR4-3200 · NVMe Gen4 · llama.cpp `0.5.0-dev` (needs `qwen4exp` arch + ngram spec support, any build after ~Sep 27 2026). All numbers from a two-round A/B-A/B bench, 2026-09-28.

```
huggingface.co/AtomicChat/Qwen3.8-Flash-Next-GGUF
  → Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64/   # 94.5 GB, 33 shards; shard 2 = 38 GB n-gram table (keep on NVMe!)
  → mmproj-Qwen3.8-Flash-Next-F16.gguf           # 0.9 GB vision encoder
```

**Daily driver — 120k context, vision included:**

```
llama-server \
  -m Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64/Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64-00001-of-00033.gguf \
  --mmproj mmproj-Qwen3.8-Flash-Next-F16.gguf --no-mmproj-offload \
  --host 127.0.0.1 --port 8014 \
  -c 122880 \
  --fit on --fit-target 128 \
  -ctk q8_0 -ctv q8_0 \
  -fa on \
  --jinja \
  -np 1 -b 4096 -ub 4096 \
  -t 20 --threads-batch 20 --prio 2 \
  --lazy-mode on \
  --spec-type ngram-mod \
  --spec-ngram-mod-n-match 60 --spec-ngram-mod-n-min 12 --spec-ngram-mod-n-max 24 \
  --cache-reuse 256 \
  --reasoning-format deepseek
```

**Fast preset:** identical command with `-c 81920`. **Full window (262k):** `-c 262144` **and** switch to `-b 2048 -ub 1024 -t 10 --threads-batch 12` (why below).

| Preset | 16k | 40k | 62k | 78k | 114k | Prefill | 
|---|---|---|---|---|---|---|
| 80k (fast) | 30.8 | 29.6 | 27.7 | 27.3 | — | 310–360 t/s | 
| 120k (daily) | 26.2 | 28.9 | — | 12–36* | 22.5 | 300–360 t/s | 
| 260k (full) | — | — | — | — | ~16 @222k+ | ~190–200 t/s | 

* decode is **draft-acceptance-bound**: `ngram-mod` drafts from the model's own n-gram table, and acceptance depends on content. I measured the *identical* rung twice at 12.1 t/s (12.11/12.06) and a different-content run at 36.3 t/s — same config, same depth. Repetitive text/code drafts well (expect the high end); hostile content can halve it. Quote bands, not points.

Compute buffers scale with `ubatch`, KV with `ctx`, and together they squeeze experts off the 16GB GPU. Every measured point fits one empirical law:

**Keep `ub × ctx` under ~540 million.** At ~671M decode collapses; at ~940M generation breaks outright.

| Your target ctx | Use | Product | Measured | 
|---|---|---|---|
| ≤ 131k | `-b 4096 -ub 4096 -t 20 -tb 20` | ≤537M | 20–32 t/s ✓ | 
| 131k–180k | `-b 2048 -ub 2048 -t 20 -tb 20` | ~410M @163k | (untested corner — should hold by the law) | 
| up to 262k | `-b 2048 -ub 1024 -t 10 -tb 12` | 268M | ~16 t/s ✓ | 
| never | `-ub 8192` | 671M even @82k | 8.9 t/s ✗ | 

Specifically: `-ub 4096` at 163k context decodes at 10.5 t/s and at 229k **fails to generate at all** — if you need the full window, the small ubatch is mandatory, and you trade ~45% of prefill speed (~200 vs ~340 t/s) for the range.

- `--lazy-mode on` + mmap (default): the 38GB n-gram shard stays on SSD; rows fault in on demand. This is what makes 64GB RAM work. Requires mmap — never combine with`--load-mode none` .
- `--no-mmproj-offload` : vision encoder on CPU =**free** ; on GPU it costs −19% decode for faster image encoding. A few extra seconds per image.
- `-t 20` on all cores including E-cores: saturates ~7–10 cores of work.**Never pin to P-cores only** — strict P-core affinity halved decode (6 physical cores + lost E-core memory streams).
- `--fit on --fit-target 128` : auto GPU/CPU layer split, beats manual`-ncmoe` by ~3%.
- CPU governor: irrelevant (memory-bandwidth-bound) — leave powersave.

- First request after load runs ~30s slow (NVMe expert page-in; 15 t/s prefill cold → 300+ warm).
- GPU util oscillates 0%↔busy ~1/sec during decode — GPU waiting on CPU-resident experts. Normal.
- ~46–50GB of resident mmap pages; under RAM pressure the kernel evicts/re-faults (dips, not crashes). In Docker, don't cap the container below ~52GB.
- VRAM steady ~15.7GB; q8_0 KV; no MTP draft head needed — the n-gram table is the drafter.

**Avoid:** `-ot 'exps=CPU'` (~4 t/s), `-ngl 99` alone (OOM), `--load-mode none` (breaks lazy-mode; needs ~93GB), EXL3/vLLM (won't fit), `-ub 8192` (see table).
