{"slug": "making-qwen-3-8-27b-fast-on-strix-halo-gfx1151", "title": "Making Qwen 3.8 27B fast on Strix Halo gfx1151", "summary": "A developer released an inference engine tailored to the Strix Halo (Bosgame) and Qwen 3.8 architecture, achieving over 550 tokens per second at 32k context depth, about twice as fast as the next closest engine. The engine, available at github.com/peonist-ai/halogen-server, includes prompt caching and is designed for agentic workflows.", "body_md": "*“They” say necessity is the mother of invention*…\n\nOverall, I’ve been pleased with my purchase of a 128GB Strix Halo (Bosgame) but the prefill has been killing me. So when Qwen3.8 27B dropped, I blew the last week or so of my life in front of Claude trying to make it usable for agentic workflows.\n\nI honestly had more success than expected.\n\nI ended up building an inference engine tailored to the Strix Halo and Qwen 3.8 architecture… It turned out being about 2x as fast at prefill as the next closest engine I could find at higher model weights. >550t/s @32k depth. There might be a little more room to go but it is approaching the wall of physics and my hardware can’t run past 2400Mhz under load. If you have good cooling and can hit closer to the 2900MHz clock, you can probably push 600+ @32k context depth.\n\nAnyway, I wasn’t going to release this, but decided if even a few people find it useful and it saves them time its worth it.\n\nI tried to make it as simple as possible to just fire up a server and run it. The instructions are in the repo. Model weights on hugging face.\n\nI’m using this with my own harness (in development) and it is definitely usable. Is it Deepseek v4 flash fast? No. But once you warm it up and you’re careful with context its been extremely useful and enjoyable to work with. The engine does prompt caching so it feels pretty good.\n\nIf you have a Strix Halo, try it out and let me know what you think. I hope it saves you time.\n\nThanks for looking.\n\nP.S - I don’t have social media or reddit so this is probably the only place I’m going to post this. If you want to leave feedback just put it in this thread, or as an issue on github.\n\nrepo is at github / peonist-ai/halogen-server", "url": "https://wpnews.pro/news/making-qwen-3-8-27b-fast-on-strix-halo-gfx1151", "canonical_source": "https://forum.level1techs.com/t/making-qwen-3-8-27b-fast-on-strix-halo-gfx1151/254422#post_1", "published_at": "2026-08-26 04:49:13+00:00", "updated_at": "2026-08-26 05:14:53.990268+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-infrastructure", "developer-tools"], "entities": ["Strix Halo", "Bosgame", "Qwen 3.8", "Claude", "Deepseek v4 flash", "peonist-ai/halogen-server"], "alternates": {"html": "https://wpnews.pro/news/making-qwen-3-8-27b-fast-on-strix-halo-gfx1151", "markdown": "https://wpnews.pro/news/making-qwen-3-8-27b-fast-on-strix-halo-gfx1151.md", "text": "https://wpnews.pro/news/making-qwen-3-8-27b-fast-on-strix-halo-gfx1151.txt", "jsonld": "https://wpnews.pro/news/making-qwen-3-8-27b-fast-on-strix-halo-gfx1151.jsonld"}}