04:49
2026-08-26
forum.level1techs.com
artificial-intelligence
Making Qwen 3.8 27B fast on Strix Halo gfx1151
A developer released an inference engine tailored to the Strix Halo (Bosgame) and Qwen 3.8 architecture, achieving over 550 tokens per second at 32k context depth, about twice as fast as the next clos…