13:12
2026-09-16
forum.level1techs.com
large-language-models
VBR k/v cache is actually usable (buun-llama)
A user running Qwen3.8-Flash-Next at IQ3 quantization on a system with 64GB of RAM, an AMD 9950X3D, a 9070XT, and a ZFS stripe of three mid-range NVMe drives reported fitting 255K tokens of context us…