10:03
2026-08-17
gist.github.com
large-language-models
Qwen3.8-27B best llama.cpp config on RTX 4090 24GB (BeeLlama, UD-Q4_K_XL, kvarn6 + kv-tail 2048 @ 130K, MTP n-max 3, fit off)
An engineer shared a quality-first llama.cpp configuration for running the Qwen3.8-27B model on a single RTX 4090 24GB GPU, achieving 60โ70 tokens per second. The setup requires the BeeLlama fork becaโฆ