Best llama.cpp config for Qwen3.8-Flash-Next (RTX 4090 24GB)
A developer has published a configuration guide for running the Qwen3.8-Flash-Next 125B MoE model with llama.cpp on an RTX 4090 24GB system, achieving up to 29 tokens per second decode speed. The setu…