Handover: Halogen + gufo behind LlamaStash on Strix Halo (Qwen3.8 Flash-Next and 27B) A developer published a handover guide for running local LLM inference on an AMD Strix Halo machine (Ryzen AI Max+ 395, gfx1151, 128 GB unified memory), configuring LlamaStash to launch Qwen3.8 Flash-Next and a 27B model through two engines: Halogen (Docker, native .hgn weights) and gufo (native build, Unsloth GGUF plus MTP head or DFlash2 draft). The guide specifies kernel 6.18.4+ with CONFIG_HSA_AMD_SVM, KFD gfx_target_version 110501 with capability bit 0x08000000, minimal UMA frame buffer, amdgpu.gttsize and ttm.pages_limit sized to installed RAM, and roughly 260 GB of disk for weights, then connects Pi, OpenCode and Claude Code to the LlamaStash proxy. You are setting up a local LLM stack on an AMD Strix Halo machine Ryzen AI Max+ 395, Radeon 8060S, gfx1151, 128 GB unified memory running native Linux. When you finish, the user can start any of these from LlamaStash TUI, CLI, or its OpenAI/Anthropic proxy , each with ready-made presets: | LlamaStash row | Engine | Weights | |---|---|---| | flash-next-halogen | Halogen Docker image | Halogen native v2 .hgn Qwen3.8-Flash-Next | | Qwen3.8-Flash-Next-UD-Q4 K XL | gufo native build | Unsloth UD-Q4 K XL GGUF + Unsloth shared MTP head | | Qwen3.8-27B-UD-Q6 K | gufo native build | Unsloth UD-Q6 K GGUF + z-lab DFlash2 draft | Then you connect Pi, OpenCode and Claude Code to the LlamaStash proxy. - Check versions first. Before you install anything, look up the newest release of each tool. If one is newer than the version below, use it, read its changelog for renamed flags or env vars, and tell the user what you changed. - LlamaStash: gh release list -R llamastash/llamastash -L 3 - gufo: gh release list -R gufo-org/gufo -L 3 - Halogen: gh api repos/peonist-ai/halogen-flash-server/commits --jq '. 0:3 .commit.message' , and CHANGELOG.md in that repo. - LlamaStash: - Versions as of 2026-10-02: LlamaStash 0.6.1, Halogen 0.16.1, gufo v0.5.0. The Halogen setup was tested on 0.16.1 with Docker 29.8. The gufo numbers were measured on gufo 0.3.0; the v0.5.0 changelog shows no change to the flags used here, but it has not been benchmarked on this setup yet. - One large model at a time. Each Flash-Next engine takes 80 to 90 GiB. - Never hard-kill a GPU server. Stop it once with SIGTERM llamastash stop , or docker stop for Halogen and wait. A server killed mid-kernel can hang the GPU and freeze the desktop. - Don't build in /tmp . It is often tmpfs. Use a directory on disk. - Ask before system-level changes: kernel command line, BIOS, groups, packages. Run these checks and report anything that fails before you continue: 1. Kernel. You need Linux 6.18.4 or later, built with CONFIG HSA AMD SVM . Halogen maps its weights through KFD's SVM and fails to register them without it. - Find the KFD node whose properties has gfx target version 110501 , under /sys/class/kfd/kfd/topology/nodes/