04:59
2026-08-16
gist.github.com
machine-learning
llama.cpp gfx1151 optimizations for Qwen3.8 27B
A developer contributed optimizations to llama.cpp for AMD's gfx1151 GPU, specifically targeting Qwen3.8 27B's SSM convolution input pattern. The patch introduces a fast LDS-transpose path for dim-0 cโฆ