cd /news/large-language-models/qwen-3-8-flash-next-config-llama-cpp Β· home β€Ί topics β€Ί large-language-models β€Ί article
[ARTICLE Β· art-141575] src=gist.github.com β†— pub= topic=large-language-models verified=true sentiment=↑ positive

Qwen 3.8 Flash Next config LLama.cpp

A developer benchmarked the Qwen3.8-Flash-Next GGUF quant (AD-4.27bpw Q4_K_M, 94.5 GB across 33 shards) on a single RTX 5060 Ti 16GB with 64GB DDR4, publishing llama.cpp server configurations that sustain 20–32 t/s decode at up to 131k context and roughly 16 t/s at 262k. The writeup derives an empirical rule that ubatch Γ— context must stay under about 540 million, above which decode collapses (8.9 t/s at 671M) and generation fails outright at ~940M, and notes that ngram-mod speculative decoding makes throughput draft-acceptance-bound, ranging from 12.1 to 36.3 t/s on identical configs.

by read3 min views1 publishedSep 28, 2026

Validated on: RTX 5060 Ti 16GB Β· i5-14500 Β· 64GB DDR4-3200 Β· NVMe Gen4 Β· llama.cpp 0.5.0-dev (needs qwen4exp arch + ngram spec support, any build after ~Sep 27 2026). All numbers from a two-round A/B-A/B bench, 2026-09-28.

huggingface.co/AtomicChat/Qwen3.8-Flash-Next-GGUF
  β†’ Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64/   # 94.5 GB, 33 shards; shard 2 = 38 GB n-gram table (keep on NVMe!)
  β†’ mmproj-Qwen3.8-Flash-Next-F16.gguf           # 0.9 GB vision encoder

Daily driver β€” 120k context, vision included:

llama-server \
  -m Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64/Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64-00001-of-00033.gguf \
  --mmproj mmproj-Qwen3.8-Flash-Next-F16.gguf --no-mmproj-offload \
  --host 127.0.0.1 --port 8014 \
  -c 122880 \
  --fit on --fit-target 128 \
  -ctk q8_0 -ctv q8_0 \
  -fa on \
  --jinja \
  -np 1 -b 4096 -ub 4096 \
  -t 20 --threads-batch 20 --prio 2 \
  --lazy-mode on \
  --spec-type ngram-mod \
  --spec-ngram-mod-n-match 60 --spec-ngram-mod-n-min 12 --spec-ngram-mod-n-max 24 \
  --cache-reuse 256 \
  --reasoning-format deepseek

Fast preset: identical command with -c 81920. Full window (262k): -c 262144 and switch to -b 2048 -ub 1024 -t 10 --threads-batch 12 (why below).

Preset 16k 40k 62k 78k 114k Prefill
80k (fast) 30.8 29.6 27.7 27.3 β€” 310–360 t/s
120k (daily) 26.2 28.9 β€” 12–36* 22.5 300–360 t/s
260k (full) β€” β€” β€” β€” ~16 @222k+ ~190–200 t/s
  • decode is draft-acceptance-bound: ngram-mod drafts from the model's own n-gram table, and acceptance depends on content. I measured the identical rung twice at 12.1 t/s (12.11/12.06) and a different-content run at 36.3 t/s β€” same config, same depth. Repetitive text/code drafts well (expect the high end); hostile content can halve it. Quote bands, not points.

Compute buffers scale with ubatch, KV with ctx, and together they squeeze experts off the 16GB GPU. Every measured point fits one empirical law:

Keep ub Γ— ctx under ~540 million. At ~671M decode collapses; at ~940M generation breaks outright.

Your target ctx Use Product Measured
≀ 131k -b 4096 -ub 4096 -t 20 -tb 20 ≀537M 20–32 t/s βœ“
131k–180k -b 2048 -ub 2048 -t 20 -tb 20 ~410M @163k (untested corner β€” should hold by the law)
up to 262k -b 2048 -ub 1024 -t 10 -tb 12 268M ~16 t/s βœ“
never -ub 8192 671M even @82k 8.9 t/s βœ—

Specifically: -ub 4096 at 163k context decodes at 10.5 t/s and at 229k fails to generate at all β€” if you need the full window, the small ubatch is mandatory, and you trade ~45% of prefill speed (~200 vs ~340 t/s) for the range.

  • --lazy-mode on + mmap (default): the 38GB n-gram shard stays on SSD; rows fault in on demand. This is what makes 64GB RAM work. Requires mmap β€” never combine with--load-mode none .

  • --no-mmproj-offload : vision encoder on CPU =free ; on GPU it costs βˆ’19% decode for faster image encoding. A few extra seconds per image.

  • -t 20 on all cores including E-cores: saturates ~7–10 cores of work.Never pin to P-cores only β€” strict P-core affinity halved decode (6 physical cores + lost E-core memory streams).

  • --fit on --fit-target 128 : auto GPU/CPU layer split, beats manual-ncmoe by ~3%.

  • CPU governor: irrelevant (memory-bandwidth-bound) β€” leave powersave.

  • First request after load runs ~30s slow (NVMe expert page-in; 15 t/s prefill cold β†’ 300+ warm).

  • GPU util oscillates 0%↔busy ~1/sec during decode β€” GPU waiting on CPU-resident experts. Normal.

  • ~46–50GB of resident mmap pages; under RAM pressure the kernel evicts/re-faults (dips, not crashes). In Docker, don't cap the container below ~52GB.

  • VRAM steady ~15.7GB; q8_0 KV; no MTP draft head needed β€” the n-gram table is the drafter.

Avoid: -ot 'exps=CPU' (~4 t/s), -ngl 99 alone (OOM), --load-mode none (breaks lazy-mode; needs ~93GB), EXL3/vLLM (won't fit), -ub 8192 (see table).

── more in #large-language-models 4 stories Β· sorted by recency
── more on @qwen 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/qwen-3-8-flash-next-…] indexed:0 read:3min 2026-09-28 Β· β€”