Validated on: RTX 5060 Ti 16GB Β· i5-14500 Β· 64GB DDR4-3200 Β· NVMe Gen4 Β· llama.cpp 0.5.0-dev (needs qwen4exp arch + ngram spec support, any build after ~Sep 27 2026). All numbers from a two-round A/B-A/B bench, 2026-09-28.
huggingface.co/AtomicChat/Qwen3.8-Flash-Next-GGUF
β Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64/ # 94.5 GB, 33 shards; shard 2 = 38 GB n-gram table (keep on NVMe!)
β mmproj-Qwen3.8-Flash-Next-F16.gguf # 0.9 GB vision encoder
Daily driver β 120k context, vision included:
llama-server \
-m Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64/Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64-00001-of-00033.gguf \
--mmproj mmproj-Qwen3.8-Flash-Next-F16.gguf --no-mmproj-offload \
--host 127.0.0.1 --port 8014 \
-c 122880 \
--fit on --fit-target 128 \
-ctk q8_0 -ctv q8_0 \
-fa on \
--jinja \
-np 1 -b 4096 -ub 4096 \
-t 20 --threads-batch 20 --prio 2 \
--lazy-mode on \
--spec-type ngram-mod \
--spec-ngram-mod-n-match 60 --spec-ngram-mod-n-min 12 --spec-ngram-mod-n-max 24 \
--cache-reuse 256 \
--reasoning-format deepseek
Fast preset: identical command with -c 81920. Full window (262k): -c 262144 and switch to -b 2048 -ub 1024 -t 10 --threads-batch 12 (why below).
| Preset | 16k | 40k | 62k | 78k | 114k | Prefill |
|---|---|---|---|---|---|---|
| 80k (fast) | 30.8 | 29.6 | 27.7 | 27.3 | β | 310β360 t/s |
| 120k (daily) | 26.2 | 28.9 | β | 12β36* | 22.5 | 300β360 t/s |
| 260k (full) | β | β | β | β | ~16 @222k+ | ~190β200 t/s |
- decode is draft-acceptance-bound:
ngram-moddrafts from the model's own n-gram table, and acceptance depends on content. I measured the identical rung twice at 12.1 t/s (12.11/12.06) and a different-content run at 36.3 t/s β same config, same depth. Repetitive text/code drafts well (expect the high end); hostile content can halve it. Quote bands, not points.
Compute buffers scale with ubatch, KV with ctx, and together they squeeze experts off the 16GB GPU. Every measured point fits one empirical law:
Keep ub Γ ctx under ~540 million. At ~671M decode collapses; at ~940M generation breaks outright.
| Your target ctx | Use | Product | Measured |
|---|---|---|---|
| β€ 131k | -b 4096 -ub 4096 -t 20 -tb 20 |
β€537M | 20β32 t/s β |
| 131kβ180k | -b 2048 -ub 2048 -t 20 -tb 20 |
~410M @163k | (untested corner β should hold by the law) |
| up to 262k | -b 2048 -ub 1024 -t 10 -tb 12 |
268M | ~16 t/s β |
| never | -ub 8192 |
671M even @82k | 8.9 t/s β |
Specifically: -ub 4096 at 163k context decodes at 10.5 t/s and at 229k fails to generate at all β if you need the full window, the small ubatch is mandatory, and you trade ~45% of prefill speed (~200 vs ~340 t/s) for the range.
-
--lazy-mode on+ mmap (default): the 38GB n-gram shard stays on SSD; rows fault in on demand. This is what makes 64GB RAM work. Requires mmap β never combine with--load-mode none. -
--no-mmproj-offload: vision encoder on CPU =free ; on GPU it costs β19% decode for faster image encoding. A few extra seconds per image. -
-t 20on all cores including E-cores: saturates ~7β10 cores of work.Never pin to P-cores only β strict P-core affinity halved decode (6 physical cores + lost E-core memory streams). -
--fit on --fit-target 128: auto GPU/CPU layer split, beats manual-ncmoeby ~3%. -
CPU governor: irrelevant (memory-bandwidth-bound) β leave powersave.
-
First request after load runs ~30s slow (NVMe expert page-in; 15 t/s prefill cold β 300+ warm).
-
GPU util oscillates 0%βbusy ~1/sec during decode β GPU waiting on CPU-resident experts. Normal.
-
~46β50GB of resident mmap pages; under RAM pressure the kernel evicts/re-faults (dips, not crashes). In Docker, don't cap the container below ~52GB.
-
VRAM steady ~15.7GB; q8_0 KV; no MTP draft head needed β the n-gram table is the drafter.
Avoid: -ot 'exps=CPU' (~4 t/s), -ngl 99 alone (OOM), --load-mode none (breaks lazy-mode; needs ~93GB), EXL3/vLLM (won't fit), -ub 8192 (see table).