Qwen 3.8 Flash Next config LLama.cpp A developer benchmarked the Qwen3.8-Flash-Next GGUF quant (AD-4.27bpw Q4_K_M, 94.5 GB across 33 shards) on a single RTX 5060 Ti 16GB with 64GB DDR4, publishing llama.cpp server configurations that sustain 20–32 t/s decode at up to 131k context and roughly 16 t/s at 262k. The writeup derives an empirical rule that ubatch × context must stay under about 540 million, above which decode collapses (8.9 t/s at 671M) and generation fails outright at ~940M, and notes that ngram-mod speculative decoding makes throughput draft-acceptance-bound, ranging from 12.1 to 36.3 t/s on identical configs. Validated on: RTX 5060 Ti 16GB · i5-14500 · 64GB DDR4-3200 · NVMe Gen4 · llama.cpp 0.5.0-dev needs qwen4exp arch + ngram spec support, any build after ~Sep 27 2026 . All numbers from a two-round A/B-A/B bench, 2026-09-28. huggingface.co/AtomicChat/Qwen3.8-Flash-Next-GGUF → Qwen3.8-Flash-Next-AD-4.27bpw-Q4 K M-M64/ 94.5 GB, 33 shards; shard 2 = 38 GB n-gram table keep on NVMe → mmproj-Qwen3.8-Flash-Next-F16.gguf 0.9 GB vision encoder Daily driver — 120k context, vision included: llama-server \ -m Qwen3.8-Flash-Next-AD-4.27bpw-Q4 K M-M64/Qwen3.8-Flash-Next-AD-4.27bpw-Q4 K M-M64-00001-of-00033.gguf \ --mmproj mmproj-Qwen3.8-Flash-Next-F16.gguf --no-mmproj-offload \ --host 127.0.0.1 --port 8014 \ -c 122880 \ --fit on --fit-target 128 \ -ctk q8 0 -ctv q8 0 \ -fa on \ --jinja \ -np 1 -b 4096 -ub 4096 \ -t 20 --threads-batch 20 --prio 2 \ --lazy-mode on \ --spec-type ngram-mod \ --spec-ngram-mod-n-match 60 --spec-ngram-mod-n-min 12 --spec-ngram-mod-n-max 24 \ --cache-reuse 256 \ --reasoning-format deepseek Fast preset: identical command with -c 81920 . Full window 262k : -c 262144 and switch to -b 2048 -ub 1024 -t 10 --threads-batch 12 why below . | Preset | 16k | 40k | 62k | 78k | 114k | Prefill | |---|---|---|---|---|---|---| | 80k fast | 30.8 | 29.6 | 27.7 | 27.3 | — | 310–360 t/s | | 120k daily | 26.2 | 28.9 | — | 12–36 | 22.5 | 300–360 t/s | | 260k full | — | — | — | — | ~16 @222k+ | ~190–200 t/s | decode is draft-acceptance-bound : ngram-mod drafts from the model's own n-gram table, and acceptance depends on content. I measured the identical rung twice at 12.1 t/s 12.11/12.06 and a different-content run at 36.3 t/s — same config, same depth. Repetitive text/code drafts well expect the high end ; hostile content can halve it. Quote bands, not points. Compute buffers scale with ubatch , KV with ctx , and together they squeeze experts off the 16GB GPU. Every measured point fits one empirical law: Keep ub × ctx under ~540 million. At ~671M decode collapses; at ~940M generation breaks outright. | Your target ctx | Use | Product | Measured | |---|---|---|---| | ≤ 131k | -b 4096 -ub 4096 -t 20 -tb 20 | ≤537M | 20–32 t/s ✓ | | 131k–180k | -b 2048 -ub 2048 -t 20 -tb 20 | ~410M @163k | untested corner — should hold by the law | | up to 262k | -b 2048 -ub 1024 -t 10 -tb 12 | 268M | ~16 t/s ✓ | | never | -ub 8192 | 671M even @82k | 8.9 t/s ✗ | Specifically: -ub 4096 at 163k context decodes at 10.5 t/s and at 229k fails to generate at all — if you need the full window, the small ubatch is mandatory, and you trade ~45% of prefill speed ~200 vs ~340 t/s for the range. - --lazy-mode on + mmap default : the 38GB n-gram shard stays on SSD; rows fault in on demand. This is what makes 64GB RAM work. Requires mmap — never combine with --load-mode none . - --no-mmproj-offload : vision encoder on CPU = free ; on GPU it costs −19% decode for faster image encoding. A few extra seconds per image. - -t 20 on all cores including E-cores: saturates ~7–10 cores of work. Never pin to P-cores only — strict P-core affinity halved decode 6 physical cores + lost E-core memory streams . - --fit on --fit-target 128 : auto GPU/CPU layer split, beats manual -ncmoe by ~3%. - CPU governor: irrelevant memory-bandwidth-bound — leave powersave. - First request after load runs ~30s slow NVMe expert page-in; 15 t/s prefill cold → 300+ warm . - GPU util oscillates 0%↔busy ~1/sec during decode — GPU waiting on CPU-resident experts. Normal. - ~46–50GB of resident mmap pages; under RAM pressure the kernel evicts/re-faults dips, not crashes . In Docker, don't cap the container below ~52GB. - VRAM steady ~15.7GB; q8 0 KV; no MTP draft head needed — the n-gram table is the drafter. Avoid: -ot 'exps=CPU' ~4 t/s , -ngl 99 alone OOM , --load-mode none breaks lazy-mode; needs ~93GB , EXL3/vLLM won't fit , -ub 8192 see table .