02:00
2026-09-09
gist.github.com
large-language-models
Qwen3.8-Flash-Next 4.05bpw EXL3 solo launch (TabbyAPI + ExLlamaV3, 5090+4090)
A developer published a serving configuration for the turboderp/Qwen3.8-Flash-Next-exl3-4.05bpw EXL3 quant, running the large MoE model split across an RTX 5090 and RTX 4090 with 192 GB of DDR5 systemβ¦