cd /news/large-language-models/qwen3-8-flash-next-4-05bpw-exl3-solo… · home › topics › large-language-models › article
[ARTICLE · art-138603] src=gist.github.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Qwen3.8-Flash-Next 4.05bpw EXL3 solo launch (TabbyAPI + ExLlamaV3, 5090+4090)

A developer published a serving configuration for the turboderp/Qwen3.8-Flash-Next-exl3-4.05bpw EXL3 quant, running the large MoE model split across an RTX 5090 and RTX 4090 with 192 GB of DDR5 system RAM via ExLlamaV3 v1.4.8 and TabbyAPI. The setup offloads 20 MoE layers to CPU and validates at 262K context with roughly 2,750 T/s prefill and 59 T/s decode, with an FP16-KV sibling config reaching about 2,370 T/s prefill and 55 T/s decode.

by read2 min views15 publishedSep 9, 2026

Serving config for the turboderp/Qwen3.8-Flash-Next-exl3-4.05bpw EXL3 quant as a single big-MoE model split across two GPUs + system RAM.

HW: RTX 5090 (GPU0) + RTX 4090 (GPU1) + 192 GB DDR5, Ryzen 9 7950X3D (16c/32t). CUDA 13.3. Engine pins: ExLlamaV3 v1.4.8 (d21d38c, dev branch), TabbyAPI 0.0.1 @ b75fe27, run via python main.py from the tabbyAPI checkout.

export CUDA_DEVICE_ORDER=PCI_BUS_ID
export CUDA_VISIBLE_DEVICES=0,1          # 0=5090, 1=4090

export EXL3_MOE_ARENA_HUGEPAGE=1         # hugepage-backed CPU expert arena
export EXL3_MOE_CPU_THREADS=24           # 24 sweet spot on 16c/32t (32 oversubscribes)
export EXL3_MOE_CPU_PIN=1                # pin to distinct physical cores
export EXL3_MOE_CPU_SWIZZLE=1            # AVX512-VBMI band-contiguous expert layout
export EXL3_MOE_CPU_STAGE_THREADS=8      # staging memcpy threads (default 4)
export EXL3_MOE_CPU_SLOT_ROWS=128        # CPU-tail rows/slot (default 64)
export EXL3_MOE_CPU_SLOTS=8              # job-ring depth (default 4)
export EXL3_MOE_STREAM_BATCH_EXPERTS=32  # experts per weight-staging DMA (default 24)

export EXL3_NGRAM_STREAM=0               # 39 GB n-gram table stays in RAM

Validated @262K: ~2,750 T/s prefill, ~59 T/s decode.

cd tabbyapi-cwd/          # CWD must hold the model-specific config.yml
python <tabbyAPI>/main.py \
  --port 8002 \
  --host 127.0.0.1 \
  --disable-auth true \
  --model-dir /path/to/llms \
  --model-name turboderp--Qwen3.8-Flash-Next-exl3-4.05bpw \
  --backend exllamav3 \
  --prompt-template qwen3.8.jinja \
  --tensor-parallel false \
  --gpu-split 27 23 \
  --cpu-moe-offload-layers 20 \
  --max-seq-len 262144 \
  --cache-size 262144 \
  --cache-mode 8,8 \
  --chunk-size 7168 \
  --output-chunking true \
  --max-batch-size 1 \
  --vision false \
  --reasoning true \
  --tool-format qwen3_coder \
  --draft-mode mtp \
  --draft-num-tokens 1 \
  --dynamic-draft false

FP16-KV sibling (slower, ~2,370/~55): identical except --cache-mode FP16 and --cpu-moe-offload-layers 22 (Q8 frees ~2.7 GiB on GPU1 so 2 MoE layers come back to GPU).

network: { host: 127.0.0.1, port: 5000, disable_auth: true }
logging: { log_prompt: false, log_generation_params: false }
model:
  model_dir: /path/to/llms
  model_name: turboderp--Qwen3.8-Flash-Next-exl3-4.05bpw
  backend: exllamav3
  tensor_parallel: false
  max_batch_size: 1
  cache_mode: FP16          # fallback only; launcher CLI overrides
  max_seq_len: -1
  chunk_size: 2048          # fallback only
  output_chunking: true
  cpu_moe_split_experts: 0
  vision: false
  reasoning: true
  tool_format: qwen3_coder
sampling:
  override_preset: defaults

Precedence: config.yml < env < CLI — the launcher passes the real split/ctx/chunk/cache as CLI, so file values are fallbacks.

force: true → overrides client-supplied sampling params.

temperature:        { override: 1.0,  force: true }
top_k:              { override: 20,   force: true }
top_p:              { override: 0.95, force: true }
min_p:              { override: 0.0,  force: true }
presence_penalty:   { override: 0.0,  force: true }
repetition_penalty: { override: 1.0,  force: true }
models:
  qwen38-flash-next-exl3-q8:
    aliases: [solo]
    proxy: "http://127.0.0.1:${PORT}"
    checkEndpoint: /v1/model
    useModelName: turboderp--Qwen3.8-Flash-Next-exl3-4.05bpw
    ttl: 0
    unloadTimeout: 30
    healthCheckTimeout: 600     # ~68 GB load, ~10 min
    concurrencyLimit: 1
    env: [ ...same env block above... ]
    cmd: /path/to/tabbyapi-qwen38-flash-next-q8-launcher.sh ${PORT}
  • --config ignores CLI--port in TabbyAPI (early return inconfig._from_args ) — pass everything as CLI args.
  • Model auto-loads on startup, so /v1/model returns 503 until the ~68 GB load completes.
  • Q8 cache-mode 8,8 quantizes only the 12 QSA layers' gathered K/V; indexer scoring stays fp16.
── more in #large-language-models 4 stories · sorted by recency
── more on @qwen3.8-flash-next 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/qwen3-8-flash-next-4…] indexed:0 read:2min 2026-09-09 · —