cd /news/large-language-models/handover-halogen-gufo-behind-llamast… · home › topics › large-language-models › article
[ARTICLE · art-144985] src=gist.github.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Handover: Halogen + gufo behind LlamaStash on Strix Halo (Qwen3.8 Flash-Next and 27B)

A developer published a handover guide for running local LLM inference on an AMD Strix Halo machine (Ryzen AI Max+ 395, gfx1151, 128 GB unified memory), configuring LlamaStash to launch Qwen3.8 Flash-Next and a 27B model through two engines: Halogen (Docker, native .hgn weights) and gufo (native build, Unsloth GGUF plus MTP head or DFlash2 draft). The guide specifies kernel 6.18.4+ with CONFIG_HSA_AMD_SVM, KFD gfx_target_version 110501 with capability bit 0x08000000, minimal UMA frame buffer, amdgpu.gttsize and ttm.pages_limit sized to installed RAM, and roughly 260 GB of disk for weights, then connects Pi, OpenCode and Claude Code to the LlamaStash proxy.

by read15 min views1 publishedOct 2, 2026

You are setting up a local LLM stack on an AMD Strix Halo machine (Ryzen AI Max+ 395, Radeon 8060S, gfx1151, 128 GB unified memory) running native Linux. When you finish, the user can start any of these from LlamaStash (TUI, CLI, or its OpenAI/Anthropic proxy), each with ready-made presets:

LlamaStash row Engine Weights
flash-next-halogen Halogen (Docker image) Halogen native v2 .hgn (Qwen3.8-Flash-Next)
Qwen3.8-Flash-Next-UD-Q4_K_XL gufo (native build) Unsloth UD-Q4_K_XL GGUF + Unsloth shared MTP head
Qwen3.8-27B-UD-Q6_K gufo (native build) Unsloth UD-Q6_K GGUF + z-lab DFlash2 draft

Then you connect Pi, OpenCode and Claude Code to the LlamaStash proxy.

  • Check versions first. Before you install anything, look up the newest release of each tool. If one is newer than the version below, use it, read its changelog for renamed flags or env vars, and tell the user what you changed.
    • LlamaStash: gh release list -R llamastash/llamastash -L 3
    • gufo: gh release list -R gufo-org/gufo -L 3
    • Halogen: gh api repos/peonist-ai/halogen-flash-server/commits --jq '.[0:3][].commit.message' , andCHANGELOG.md in that repo.
  • LlamaStash:
  • Versions as of 2026-10-02: LlamaStash 0.6.1, Halogen 0.16.1, gufo v0.5.0. The Halogen setup was tested on 0.16.1 with Docker 29.8. The gufo numbers were measured on gufo 0.3.0; the v0.5.0 changelog shows no change to the flags used here, but it has not been benchmarked on this setup yet.
  • One large model at a time. Each Flash-Next engine takes 80 to 90 GiB.
  • Never hard-kill a GPU server. Stop it once with SIGTERM (llamastash stop , ordocker stop for Halogen) and wait. A server killed mid-kernel can hang the GPU and freeze the desktop.
  • Don't build in /tmp. It is often tmpfs. Use a directory on disk.
  • Ask before system-level changes: kernel command line, BIOS, groups, packages.

Run these checks and report anything that fails before you continue:

  1. Kernel. You need Linux 6.18.4 or later, built withCONFIG_HSA_AMD_SVM . Halogen maps its weights through KFD's SVM and fails to register them without it.
  • Find the KFD node whose properties hasgfx_target_version 110501 , under/sys/class/kfd/kfd/topology/nodes/<n>/properties .
  • Its capability must have bit0x08000000 set.
  1. Find the KFD node whose
  2. BIOS. The UMA frame buffer ("dedicated graphics memory") must be at its smallest explicit value (for example 512 MB), not Auto. Halogen wants the memory as normal RAM.
  3. Kernel command line (cat /proc/cmdline ).
  • Set amdgpu.gttsize andttm.pages_limit to about the installed RAM. For 128 GB that isamdgpu.gttsize=126976 ttm.pages_limit=32505856 ; the Halogen README has the rows for 96 and 64 GB.
  • Optional:amd_iommu=off makes Halogen prefill 13 to 16% faster (Halogen's own measurement). It also disables the NPU and can break USB4 docks. Ask the user before adding it.
  1. Set
  2. Groups. The user must be in thevideo andrender groups, and needs read/write access to/dev/kfd and/dev/dri/renderD* .
  3. Tools: Docker,socat ,curl ,setpriv (util-linux),uv ,git ,gh ,cmake 3.21+,ninja ,pkg-config , a C++20 compiler, and the dev packages for ICU, libcurl, OpenSSL, libpng, libjpeg and libwebp, plusffmpeg .
  4. Disk: about 260 GB free for the weights.
  • Halogen: about 118 GB.
  • Flash-Next GGUF + MTP head: about 114 GB.
  • 27B + draft: about 24 GB.
  1. Power. On AC power, with/sys/firmware/acpi/platform_profile atperformance if the machine has it. Battery or a low profile can halve the speed.
curl -fsSL https://llamastash.dev/install.sh | sh     # or: yay -S llamastash / brew / cargo install llamastash
llamastash --version
llamastash init        # first run only: installs llama-server, writes the config
  • You need 0.6.1 or later. It added the generic-entry fields that give Halogen its thinking-effort controls in Pi and OpenCode (step 6). install.sh installs the latest GitHub release. The AUR, Homebrew and crates.io packages can trail a release by a few hours, so checkllamastash --version .
  • The config is at ~/.config/llamastash/config.yaml .
  • LlamaStash scans ~/.cache/huggingface/hub for GGUFs. If the weights go anywhere else (a customHF_HOME ), add thathub directory tomodel_paths: in the config.

Use llamastash pull. Files land in the standard HF cache ($HF_HOME/hub, by default ~/.cache/huggingface/hub), which LlamaStash scans.

for f in qwen38-flash-next-v2.hgn qwen38-flash-next-ngram.hgn \
         tokenizer/chat_template.jinja tokenizer/generation_config.json tokenizer/merges.txt \
         tokenizer/tokenizer.json tokenizer/tokenizer_config.json tokenizer/vocab.json; do
  llamastash pull --no-companions --json "peonist-ai/halogen-qwen3.8-flash-next:$f" | jq -r .revision
done

llamastash pull --no-companions --json unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf | jq -r .revision
llamastash pull --no-companions --json unsloth/Qwen3.8-Flash-Next-GGUF:MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf | jq -r .revision

llamastash pull --no-companions --json unsloth/Qwen3.8-27B-GGUF:Qwen3.8-27B-UD-Q6_K.gguf | jq -r .revision
llamastash pull --no-companions --json z-lab/Qwen3.8-27B-DFlash2-GGUF:Qwen3.8-27B-DFlash2-Q4_K_M.gguf | jq -r .revision

llamastash pull --no-companions --json unsloth/Qwen3.8-Flash-Next-GGUF:mmproj-BF16.gguf | jq -r .revision
llamastash pull --no-companions --json unsloth/Qwen3.8-27B-GGUF:mmproj-BF16.gguf | jq -r .revision

Always pass --no-companions. Without it, pull also fetches one mmproj projector and one MTP head per repo, and it picks the wrong ones for gufo. For Flash-Next it would getmtp-Qwen3.8-Flash-Next-BF16.gguf , not theshared-Q8_0 head. For the projector it would getmmproj-F16.gguf . #

gufo only reads mmproj-BF16.gguf. It looks in the model file's folder and one folder up, which is the snapshot root for both repos. With onlymmproj-F16.gguf , an image request fails with "image input requires a matching --mmproj BF16 sidecar". #

pull has no revision flag , so it fetches each repo's currentmain . The printedrevision is thesnapshots/<revision> directory in the paths of steps 4 and 6. As of 2026-10-02 they match the ones written there:

  • Halogen fd91981e...
  • Flash-Next 38bb39ee...
  • 27B 4ca72078...
  • DFlash2 2d9571f8... If one differs, use the printed value.
  • Halogen

Non-GGUF files:pull 's docs only mention.gguf file pins. Pinning other files (the.hgn and tokenizer files) works; it was tested on LlamaStash 0.6.0 with a tokenizer file. #

Snapshot files are symlinks into hub/blobs/ . Anything that reads them from a container must mount the wholehub/ directory.

docker pull ghcr.io/peonist-ai/halogen-flash-server:0.16.1

Save the wrapper below as ~/.local/bin/halogen-serve and chmod +x it. Set HUB to the absolute path of your hub/ directory.

#!/bin/sh
port="$1"
name="llamastash-halogen-$port"
net=halogen-net
HUB=/home/USER/.cache/huggingface/hub    # whole hub dir: snapshot files are symlinks into blobs/
REV=fd91981e3e8117ddbb324fb7efb4c1d6df8fe9e2
IMAGE=ghcr.io/peonist-ai/halogen-flash-server:0.16.1
hg=/hub/models--peonist-ai--halogen-qwen3.8-flash-next/snapshots/$REV
docker rm -f "$name" >/dev/null 2>&1
docker network inspect "$net" >/dev/null 2>&1 || docker network create --internal "$net" >/dev/null || exit 1
docker run -d --rm --name "$name" --network "$net" \
  --device /dev/kfd --device /dev/dri \
  --group-add "$(getent group video | cut -d: -f3)" --group-add "$(getent group render | cut -d: -f3)" \
  --ipc=host --ulimit memlock=-1:-1 -v "$HUB":/hub:ro \
  -e HALOGEN_API_PORT=8080 -e HALOGEN_MODEL_ID \
  -e HALOGEN_CHECKPOINT="$hg/qwen38-flash-next-v2.hgn" \
  -e HALOGEN_TOKENIZER="$hg/tokenizer" \
  -e HALOGEN_CTX -e HALOGEN_KV_POOL_POSITIONS -e HALOGEN_MTP_DEPTH -e HALOGEN_REASONING_EFFORT \
  -e HALOGEN_MAX_TOKENS_DEFAULT=16384 -e HALOGEN_TEMPERATURE -e HALOGEN_TOP_P=0.95 -e HALOGEN_TOP_K=20 \
  "$IMAGE" >/dev/null || exit 1
trap 'docker stop -t 60 "$name" >/dev/null 2>&1' TERM INT
docker logs -f "$name" 2>&1 &
ip=$(docker inspect -f "{{(index .NetworkSettings.Networks \"$net\").IPAddress}}" "$name")
until curl -so /dev/null "http://$ip:8080/v1/models"; do
  [ "$(docker inspect -f '{{.State.Running}}' "$name" 2>/dev/null)" = true ] || exit 1
  sleep 2
done
setpriv --pdeathsig TERM socat TCP-LISTEN:"$port",bind=127.0.0.1,reuseaddr,fork TCP:"$ip":8080 &
fwd=$!
docker wait "$name" >/dev/null &
wait $!
kill "$fwd" 2>/dev/null
  • Halogen finds qwen38-flash-next-ngram.hgn next to the checkpoint by itself. Do not setHALOGEN_MTP_HEAD with v2, because the draft head is inside the checkpoint.
  • No outbound network. On the--internal network the container can't resolve DNS or reach any outside address.HF_HUB_OFFLINE=1 is set in the image, andHALOGEN_DOWNLOAD stays unset, so it never needs to. To check:docker exec <container> python3 -c 'import urllib.request as u; u.urlopen("https://example.com", timeout=5)' must fail with "Temporary failure in name resolution".
  • The container can still connect to host services that listen on all interfaces (ss -ltn lists them as0.0.0.0 ).
  • Podman: not tested. Replace the two--group-add flags with--group-add keep-groups (rootless Podman needscrun for that). Rootless Podman keeps containers in its own network namespace, so the host may not reach the container IP. If the socat forward fails, drop the internal network and publish with-p "127.0.0.1:$port:8080" ; outbound traffic is then open.

This is the setup the numbers were measured on. It uses the ROCm 10 pip SDK in its own venv, so it doesn't need a system ROCm. LLMS is any directory on disk.

LLMS=$HOME/llms; SDK=$LLMS/rocm10-sdk; mkdir -p $SDK/tmp
git clone https://github.com/gufo-org/gufo $LLMS/gufo && git -C $LLMS/gufo checkout v0.5.0

uv venv --seed -p 3.12 $SDK/.venv
TMPDIR=$SDK/tmp PIP_NO_CACHE_DIR=1 $SDK/.venv/bin/python -m pip install \
  --index-url https://stable.repo.amd.com/rocm/whl-next/ "rocm[libraries,devel,device-gfx1151]==10.0.0"
$SDK/.venv/bin/rocm-sdk init

R=$($SDK/.venv/bin/rocm-sdk path --root)
export PATH="$R/bin:$PATH" HIP_PATH="$R" ROCM_PATH="$R"
cd $LLMS/gufo
cmake --preset release -B build/release-rocm10 \
  -DCMAKE_PREFIX_PATH="$R" -DCMAKE_HIP_COMPILER="$R/lib/llvm/bin/clang++" \
  "-DCMAKE_BUILD_RPATH=$R/lib;$R/lib/rocm_sysdeps/lib;$R/lib/llvm/lib"
cmake --build build/release-rocm10 -j 16
./build/release-rocm10/gufo diagnose
  • GCC 16 or newer (Arch, for example) breaks ROCm clang HIP compiles in<format> . Install GCC 15 and add the following to thecmake --preset line:
-DCMAKE_C_COMPILER=gcc-15 -DCMAKE_CXX_COMPILER=g++-15 -DCMAKE_HIP_FLAGS=--gcc-install-dir=$(dirname $(gcc-15 -print-libgcc-file-name))
  • The RPATH points the binary at the SDK's libraries, so it runs without any env vars.
  • Fallback: gufo's own qualified toolchain is a system ROCm 7.2.3 with GCC 15.3. Follow "Build from source / Without Nix" in gufo's README. On the reference machine, ROCm 7.2.4 and ROCm 10 builds were within 3% of each other.
  1. Stop the daemon first: llamastash daemon stop . The config is read only at start.
  2. Merge the YAML below into ~/.config/llamastash/config.yaml , keeping anything already there (for examplebackend.llamacpp ).
  3. Replace every /ABS/... path with a real absolute path:
  • /ABS/llms =$LLMS from step 5
  • /ABS/hub = thehub/ directory from step 3
  1. Start the daemon again: llamastash daemon start .

Presets set only the context window, plus the KV pool for Halogen. Thinking effort is not a preset. Each engine defaults to xhigh, and a client that sends reasoning_effort (such as Pi's thinking level or Claude Code's /effort) overrides it per request.

proxy:
  idle_ttl_secs: 0 # a loaded model stays until you stop it
backend:
  generic:
    servers:
      - name: gufo
        model: Qwen3.8-Flash-Next-UD-Q4_K_XL
        binary: /ABS/llms/gufo/build/release-rocm10/gufo
        args:
          [
            serve,
            --host,
            "{host}",
            --port,
            "{port}",
            --sessions,
            "1",
            llm,
            --served-model-name,
            "{name}",
            --model,
            "{model}",
            --mtp-model,
            /ABS/hub/models--unsloth--Qwen3.8-Flash-Next-GGUF/snapshots/38bb39ee97821de2c9009abb7e93950eec396e66/MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf,
            --top-p,
            "0.95",
            --top-k,
            "20",
            --reasoning-effort,
            xhigh,
          ]
        knobs:
          - { flag: --context, ctx: true, default: "131072" }
          - { flag: --speculative, default: mtp }
          - { flag: --temperature, default: "1.0" }
          - --seed
        ready: /ready
        stop_grace_secs: 60
        rewrite_model: true
      - name: gufo-27b
        model: Qwen3.8-27B-UD-Q6_K
        binary: /ABS/llms/gufo/build/release-rocm10/gufo
        args:
          [
            serve,
            --host,
            "{host}",
            --port,
            "{port}",
            --sessions,
            "1",
            llm,
            --served-model-name,
            "{name}",
            --model,
            "{model}",
            --speculative,
            dflash2,
            --dflash-model,
            /ABS/hub/models--z-lab--Qwen3.8-27B-DFlash2-GGUF/snapshots/2d9571f8ce46e151f61c6499c99dee6079e1d610/Qwen3.8-27B-DFlash2-Q4_K_M.gguf,
            --top-p,
            "0.95",
            --top-k,
            "20",
            --reasoning-effort,
            xhigh,
          ]
        knobs:
          - { flag: --context, ctx: true, default: "131072" }
          - { flag: --temperature, default: "1.0" }
          - --seed
        ready: /ready
        stop_grace_secs: 60
        rewrite_model: true
      - name: flash-next-halogen
        binary: ~/.local/bin/halogen-serve
        args: ["{port}"]
        arch: qwen4exp
        params: 177B
        quant: v2-4.16bpw
        ctx: 262144
        reasoning_effort: [none, low, medium, xhigh]
        reasoning_effort_default: xhigh # same as halogen-effort's default
        memory_gib: 87
        knobs:
          - {
              flag: --ctx-window,
              id: halogen-ctx,
              ctx: true,
              default: "131072",
            }
          - { flag: --halogen-kv-pool, id: halogen-pool, default: "131072" }
          - { flag: --halogen-temperature, id: halogen-temp, default: "1.0" }
          - { flag: --halogen-mtp-depth, id: halogen-mtp-depth, default: "3" }
          - {
              flag: --halogen-reasoning-effort,
              id: halogen-effort,
              default: xhigh,
            }
        env:
          HALOGEN_MODEL_ID: "{name}"
          HALOGEN_CTX: "{halogen-ctx}"
          HALOGEN_KV_POOL_POSITIONS: "{halogen-pool}"
          HALOGEN_TEMPERATURE: "{halogen-temp}"
          HALOGEN_MTP_DEPTH: "{halogen-mtp-depth}"
          HALOGEN_REASONING_EFFORT: "{halogen-effort}"
        ready: /v1/models
        stop_grace_secs: 90
        ready_timeout_secs: 600
presets:
  Qwen3.8-27B-UD-Q6_K.gguf:
    default: 256k
    entries:
      128k:
        server: generic-gufo-27b
        knobs: { context: 131072, temperature: "1.0" }
      256k:
        server: generic-gufo-27b
        knobs: { context: 262144, temperature: "1.0" }
  Qwen3.8-Flash-Next-*:
    default: 128k
    entries:
      128k:
        server: generic-gufo
        knobs: { context: 131072, speculative: mtp, temperature: "1.0" }
      256k:
        server: generic-gufo
        knobs: { context: 262144, speculative: mtp, temperature: "1.0" }
  flash-next-halogen:
    default: solo-256k
    entries:
      solo-256k:
        knobs: { halogen-ctx: 262144, halogen-pool: 262144 }
      dual-128k:
        knobs: { halogen-ctx: 131072, halogen-pool: 262144 }
      triple-128k:
        knobs: { halogen-ctx: 131072, halogen-pool: 393216 }
      dual-256k:
        knobs: { halogen-ctx: 262144, halogen-pool: 524288 }
      triple-256k:
        knobs: { halogen-ctx: 262144, halogen-pool: 786432 }
  1. Run llamastash list . You should see:
  • the two GGUF rows with backend llamacpp|generic
  • a flash-next-halogen row with ctx 262144 and size 87G
  • llamastash presets list <model> showing the presets above
  1. the two GGUF rows with backend
  2. Test each engine one at a time. Start it, send one chat request through the proxy, then stop it:
  • llamastash start flash-next-halogen --preset solo-256k
  • llamastash start Qwen3.8-Flash-Next-UD-Q4_K_XL --preset 128k
  • llamastash start Qwen3.8-27B-UD-Q6_K --preset 128k
  • The proxy is at http://$(llamastash status --json | jq -r .proxy.listen)/v1 . The model id is<id>@<preset> , as listed byGET /v1/models .
  • Stop with llamastash stop <model> . Wait for the stop to finish before starting the next engine.
  1. Watch llamastash logs <model> . Expected:
  • Halogen cold load takes 85 to 105 s, or about 6 s when the weights are still in page cache. gufo takes about 22 s.
  • Halogen's startup log prints a memory line. If it warns about compaction stalls , runecho 1 | sudo tee /proc/sys/vm/compact_memory before the next start. Its log suggests this instead of a reboot.
  • After any unclean exit, check cat /sys/class/drm/card*/device/mem_info_gtt_used . Tens of GiB used with nothing running means the GPU kept the memory, and only a reboot frees it.
  1. Reference numbers from an ASUS ROG Flow Z13 (Strix Halo 128 GB) at a 70 W TDP, 64k window with a 32k prompt, temperature 1.0 / top_p 0.95 / top_k 20. A desktop at 85 to 100 W should be faster.
Engine Prefill t/s Decode t/s Draft accept
Halogen 0.15.1, v2 1,191 39.4 85%
gufo 0.3.0, Flash-Next, ROCm 10 1,042 34.8 75%
gufo, 27B UD-Q6_K + DFlash2 405 17 to 21 46%

Halogen 0.16.1 on the same laptop (performance profile, on AC), with a 30k-token wikitext prompt at the 256k window: prefill 1,276 to 1,362 t/s, decode 42.3 to 43.4 t/s, draft accept 78 to 81%. The prompt differs from the table's, so this is not a direct comparison.

LlamaStash's integrations command registers every favorite model, one entry per preset, as <id>@<preset>. So mark the three rows as favorites first:

llamastash favorites add Qwen3.8-Flash-Next-UD-Q4_K_XL
llamastash favorites add Qwen3.8-27B-UD-Q6_K
llamastash favorites add flash-next-halogen
llamastash integrations pi opencode claude-code
  • Pi: patches~/.pi/agent/models.json with allamastash provider. The API key resolves through!llamastash api-key . Pick a model likeflash-next-halogen@solo-256k in Pi.
  • OpenCode: patches its config with the proxy URL and the same model list.
  • Claude Code: writes~/.config/llamastash/claude-code.sh and doesn't touch~/.claude/settings.json . Run it withsource ~/.config/llamastash/claude-code.sh && claude .
    • Use Halogen for Claude Code. Set bothANTHROPIC_MODEL andANTHROPIC_SMALL_FAST_MODEL in that file to a Halogen preset id, such asflash-next-halogen@solo-256k .
    • gufo can't serve Claude Code: as of v0.5.0 its /v1/messages rejects tools,thinking ,output_config and streaming.
  • Use
  • Add codex ,aider ,continue ,zed orenv-sh to the same command for other tools.
  • Re-run llamastash integrations ... after you add favorites or presets.
  • Effort and image fields.integrations writes them for every row:reasoning , plusthinkingLevelMap in Pi orvariants in OpenCode, with levels off (none ), low, medium and xhigh.
    • The gufo rows get theirs from the GGUF chat template, and image input from the mmproj file in step 3.
    • The Halogen row gets effort from the reasoning_effort fields in step 6 and no image input.
    • To check, open Pi's /model and pick aflash-next-halogen@... entry. Its thinking-level control should list off, low, medium and xhigh.
    • gufo needs mmproj-BF16.gguf for images. The integration only checks that some projector file is present, so a row can list image input while gufo still refuses images.
  • The gufo rows get theirs from the GGUF chat template, and image input from the

Tell the user:

  • the versions installed
  • anything you changed from this guide, and why
  • any host check that failed
  • the prefill and decode numbers from step 7
── more in #large-language-models 4 stories · sorted by recency
── more on @llamastash 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/handover-halogen-guf…] indexed:0 read:15min 2026-10-02 · —