{"slug": "handover-halogen-gufo-behind-llamastash-on-strix-halo-qwen3-8-flash-next-and-27b", "title": "Handover: Halogen + gufo behind LlamaStash on Strix Halo (Qwen3.8 Flash-Next and 27B)", "summary": "A developer published a handover guide for running local LLM inference on an AMD Strix Halo machine (Ryzen AI Max+ 395, gfx1151, 128 GB unified memory), configuring LlamaStash to launch Qwen3.8 Flash-Next and a 27B model through two engines: Halogen (Docker, native .hgn weights) and gufo (native build, Unsloth GGUF plus MTP head or DFlash2 draft). The guide specifies kernel 6.18.4+ with CONFIG_HSA_AMD_SVM, KFD gfx_target_version 110501 with capability bit 0x08000000, minimal UMA frame buffer, amdgpu.gttsize and ttm.pages_limit sized to installed RAM, and roughly 260 GB of disk for weights, then connects Pi, OpenCode and Claude Code to the LlamaStash proxy.", "body_md": "You are setting up a local LLM stack on an AMD Strix Halo machine (Ryzen AI Max+ 395, Radeon 8060S, gfx1151, 128 GB unified memory) running native Linux. When you finish, the user can start any of these from LlamaStash (TUI, CLI, or its OpenAI/Anthropic proxy), each with ready-made presets:\n\n| LlamaStash row | Engine | Weights | \n|---|---|---|\n| `flash-next-halogen` | Halogen (Docker image) | Halogen native v2 `.hgn` (Qwen3.8-Flash-Next) | \n| `Qwen3.8-Flash-Next-UD-Q4_K_XL` | gufo (native build) | Unsloth UD-Q4_K_XL GGUF + Unsloth shared MTP head | \n| `Qwen3.8-27B-UD-Q6_K` | gufo (native build) | Unsloth UD-Q6_K GGUF + z-lab DFlash2 draft | \n\nThen you connect Pi, OpenCode and Claude Code to the LlamaStash proxy.\n\n- **Check versions first.** Before you install anything, look up the newest release of each tool. If one is newer than the version below, use it, read its changelog for renamed flags or env vars, and tell the user what you changed.\n  - LlamaStash: `gh release list -R llamastash/llamastash -L 3`\n  - gufo: `gh release list -R gufo-org/gufo -L 3`\n  - Halogen: `gh api repos/peonist-ai/halogen-flash-server/commits --jq '.[0:3][].commit.message'` , and`CHANGELOG.md` in that repo.\n- LlamaStash: \n- **Versions as of 2026-10-02:** LlamaStash 0.6.1, Halogen 0.16.1, gufo v0.5.0. The Halogen setup was tested on 0.16.1 with Docker 29.8. The gufo numbers were measured on gufo 0.3.0; the v0.5.0 changelog shows no change to the flags used here, but it has not been benchmarked on this setup yet.\n- **One large model at a time.** Each Flash-Next engine takes 80 to 90 GiB.\n- **Never hard-kill a GPU server.** Stop it once with SIGTERM (`llamastash stop` , or`docker stop` for Halogen) and wait. A server killed mid-kernel can hang the GPU and freeze the desktop.\n- **Don't build in `/tmp`.** It is often tmpfs. Use a directory on disk.\n- **Ask before system-level changes:** kernel command line, BIOS, groups, packages.\n\nRun these checks and report anything that fails before you continue:\n\n1. **Kernel.** You need Linux 6.18.4 or later, built with`CONFIG_HSA_AMD_SVM` . Halogen maps its weights through KFD's SVM and fails to register them without it.\n  - Find the KFD node whose `properties` has`gfx_target_version 110501` , under`/sys/class/kfd/kfd/topology/nodes/<n>/properties` .\n  - Its `capability` must have bit`0x08000000` set.\n2. Find the KFD node whose \n3. **BIOS.** The UMA frame buffer (\"dedicated graphics memory\") must be at its smallest explicit value (for example 512 MB), not Auto. Halogen wants the memory as normal RAM.\n4. **Kernel command line** (`cat /proc/cmdline` ).\n  - Set `amdgpu.gttsize` and`ttm.pages_limit` to about the installed RAM. For 128 GB that is`amdgpu.gttsize=126976 ttm.pages_limit=32505856` ; the Halogen README has the rows for 96 and 64 GB.\n  - **Optional:**`amd_iommu=off` makes Halogen prefill 13 to 16% faster (Halogen's own measurement). It also disables the NPU and can break USB4 docks. Ask the user before adding it.\n5. Set \n6. **Groups.** The user must be in the`video` and`render` groups, and needs read/write access to`/dev/kfd` and`/dev/dri/renderD*` .\n7. **Tools:** Docker,`socat` ,`curl` ,`setpriv` (util-linux),`uv` ,`git` ,`gh` ,`cmake` 3.21+,`ninja` ,`pkg-config` , a C++20 compiler, and the dev packages for ICU, libcurl, OpenSSL, libpng, libjpeg and libwebp, plus`ffmpeg` .\n8. **Disk:** about 260 GB free for the weights.\n  - Halogen: about 118 GB.\n  - Flash-Next GGUF + MTP head: about 114 GB.\n  - 27B + draft: about 24 GB.\n9. **Power.** On AC power, with`/sys/firmware/acpi/platform_profile` at`performance` if the machine has it. Battery or a low profile can halve the speed.\n\n```\ncurl -fsSL https://llamastash.dev/install.sh | sh     # or: yay -S llamastash / brew / cargo install llamastash\nllamastash --version\nllamastash init        # first run only: installs llama-server, writes the config\n```\n\n- You need 0.6.1 or later. It added the generic-entry fields that give Halogen its thinking-effort controls in Pi and OpenCode (step 6). `install.sh` installs the latest GitHub release. The AUR, Homebrew and crates.io packages can trail a release by a few hours, so check`llamastash --version` .\n- The config is at `~/.config/llamastash/config.yaml` .\n- LlamaStash scans `~/.cache/huggingface/hub` for GGUFs. If the weights go anywhere else (a custom`HF_HOME` ), add that`hub` directory to`model_paths:` in the config.\n\nUse `llamastash pull`. Files land in the standard HF cache (`$HF_HOME/hub`, by default `~/.cache/huggingface/hub`), which LlamaStash scans.\n\n```\n# Halogen native weights: v2 checkpoint (draft head inside), its n-gram table, tokenizer\nfor f in qwen38-flash-next-v2.hgn qwen38-flash-next-ngram.hgn \\\n         tokenizer/chat_template.jinja tokenizer/generation_config.json tokenizer/merges.txt \\\n         tokenizer/tokenizer.json tokenizer/tokenizer_config.json tokenizer/vocab.json; do\n  llamastash pull --no-companions --json \"peonist-ai/halogen-qwen3.8-flash-next:$f\" | jq -r .revision\ndone\n\n# Flash-Next GGUF for gufo: pinning shard 1 pulls all 4 shards. Then the shared MTP head.\nllamastash pull --no-companions --json unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf | jq -r .revision\nllamastash pull --no-companions --json unsloth/Qwen3.8-Flash-Next-GGUF:MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf | jq -r .revision\n\n# 27B for gufo + DFlash2 draft\nllamastash pull --no-companions --json unsloth/Qwen3.8-27B-GGUF:Qwen3.8-27B-UD-Q6_K.gguf | jq -r .revision\nllamastash pull --no-companions --json z-lab/Qwen3.8-27B-DFlash2-GGUF:Qwen3.8-27B-DFlash2-Q4_K_M.gguf | jq -r .revision\n\n# Vision projectors, so gufo accepts images (about 0.9 GB each)\nllamastash pull --no-companions --json unsloth/Qwen3.8-Flash-Next-GGUF:mmproj-BF16.gguf | jq -r .revision\nllamastash pull --no-companions --json unsloth/Qwen3.8-27B-GGUF:mmproj-BF16.gguf | jq -r .revision\n```\n\n- \n**Always pass `--no-companions`.** Without it,` pull` also fetches one mmproj projector and one MTP head per repo, and it picks the wrong ones for gufo. For Flash-Next it would get`mtp-Qwen3.8-Flash-Next-BF16.gguf` , not the`shared-Q8_0` head. For the projector it would get`mmproj-F16.gguf` .\n- \n**gufo only reads `mmproj-BF16.gguf`.** It looks in the model file's folder and one folder up, which is the snapshot root for both repos. With only`mmproj-F16.gguf` , an image request fails with \"image input requires a matching --mmproj BF16 sidecar\".\n- \n**`pull` has no revision flag** , so it fetches each repo's current`main` . The printed`revision` is the`snapshots/<revision>` directory in the paths of steps 4 and 6. As of 2026-10-02 they match the ones written there:\n  - Halogen `fd91981e...`\n  - Flash-Next `38bb39ee...`\n  - 27B `4ca72078...`\n  - DFlash2 `2d9571f8...`\n If one differs, use the printed value.\n- Halogen \n- \n**Non-GGUF files:**`pull` 's docs only mention`.gguf` file pins. Pinning other files (the`.hgn` and tokenizer files) works; it was tested on LlamaStash 0.6.0 with a tokenizer file.\n- \nSnapshot files are symlinks into `hub/blobs/` . Anything that reads them from a container must mount the whole`hub/` directory.\n\n```\ndocker pull ghcr.io/peonist-ai/halogen-flash-server:0.16.1\n```\n\nSave the wrapper below as `~/.local/bin/halogen-serve` and `chmod +x` it. Set `HUB` to the absolute path of your `hub/` directory.\n\n``` bash\n#!/bin/sh\n# LlamaStash generic wrapper for Halogen (Qwen3.8-Flash-Next, native v2 .hgn).\n# $1 = port; per-launch HALOGEN_* values come from the LlamaStash entry env.\n# The server is closed source, so it runs on an internal network with no\n# outbound access. Docker can't publish ports from one, so socat forwards\n# loopback to the container; pdeathsig stops socat if this script dies.\nport=\"$1\"\nname=\"llamastash-halogen-$port\"\nnet=halogen-net\nHUB=/home/USER/.cache/huggingface/hub    # whole hub dir: snapshot files are symlinks into blobs/\nREV=fd91981e3e8117ddbb324fb7efb4c1d6df8fe9e2\nIMAGE=ghcr.io/peonist-ai/halogen-flash-server:0.16.1\nhg=/hub/models--peonist-ai--halogen-qwen3.8-flash-next/snapshots/$REV\ndocker rm -f \"$name\" >/dev/null 2>&1\ndocker network inspect \"$net\" >/dev/null 2>&1 || docker network create --internal \"$net\" >/dev/null || exit 1\ndocker run -d --rm --name \"$name\" --network \"$net\" \\\n  --device /dev/kfd --device /dev/dri \\\n  --group-add \"$(getent group video | cut -d: -f3)\" --group-add \"$(getent group render | cut -d: -f3)\" \\\n  --ipc=host --ulimit memlock=-1:-1 -v \"$HUB\":/hub:ro \\\n  -e HALOGEN_API_PORT=8080 -e HALOGEN_MODEL_ID \\\n  -e HALOGEN_CHECKPOINT=\"$hg/qwen38-flash-next-v2.hgn\" \\\n  -e HALOGEN_TOKENIZER=\"$hg/tokenizer\" \\\n  -e HALOGEN_CTX -e HALOGEN_KV_POOL_POSITIONS -e HALOGEN_MTP_DEPTH -e HALOGEN_REASONING_EFFORT \\\n  -e HALOGEN_MAX_TOKENS_DEFAULT=16384 -e HALOGEN_TEMPERATURE -e HALOGEN_TOP_P=0.95 -e HALOGEN_TOP_K=20 \\\n  \"$IMAGE\" >/dev/null || exit 1\n# One clean stop: SIGTERM becomes `docker stop`; the image SIGKILLs its engine\n# 30 s after that, so -t stays above 30 and stop_grace_secs above -t.\ntrap 'docker stop -t 60 \"$name\" >/dev/null 2>&1' TERM INT\ndocker logs -f \"$name\" 2>&1 &\n# Forward only once the API listens; earlier, socat logs every readiness probe as refused.\nip=$(docker inspect -f \"{{(index .NetworkSettings.Networks \\\"$net\\\").IPAddress}}\" \"$name\")\nuntil curl -so /dev/null \"http://$ip:8080/v1/models\"; do\n  [ \"$(docker inspect -f '{{.State.Running}}' \"$name\" 2>/dev/null)\" = true ] || exit 1\n  sleep 2\ndone\nsetpriv --pdeathsig TERM socat TCP-LISTEN:\"$port\",bind=127.0.0.1,reuseaddr,fork TCP:\"$ip\":8080 &\nfwd=$!\ndocker wait \"$name\" >/dev/null &\nwait $!\nkill \"$fwd\" 2>/dev/null\n```\n\n- Halogen finds `qwen38-flash-next-ngram.hgn` next to the checkpoint by itself. Do not set`HALOGEN_MTP_HEAD` with v2, because the draft head is inside the checkpoint.\n- **No outbound network.** On the`--internal` network the container can't resolve DNS or reach any outside address.`HF_HUB_OFFLINE=1` is set in the image, and`HALOGEN_DOWNLOAD` stays unset, so it never needs to. To check:`docker exec <container> python3 -c 'import urllib.request as u; u.urlopen(\"https://example.com\", timeout=5)'` must fail with \"Temporary failure in name resolution\".\n- The container can still connect to host services that listen on all interfaces (`ss -ltn` lists them as`0.0.0.0` ).\n- **Podman:** not tested. Replace the two`--group-add` flags with`--group-add keep-groups` (rootless Podman needs`crun` for that). Rootless Podman keeps containers in its own network namespace, so the host may not reach the container IP. If the socat forward fails, drop the internal network and publish with`-p \"127.0.0.1:$port:8080\"` ; outbound traffic is then open.\n\nThis is the setup the numbers were measured on. It uses the ROCm 10 pip SDK in its own venv, so it doesn't need a system ROCm. `LLMS` is any directory on disk.\n\n```\nLLMS=$HOME/llms; SDK=$LLMS/rocm10-sdk; mkdir -p $SDK/tmp\ngit clone https://github.com/gufo-org/gufo $LLMS/gufo && git -C $LLMS/gufo checkout v0.5.0\n\nuv venv --seed -p 3.12 $SDK/.venv\nTMPDIR=$SDK/tmp PIP_NO_CACHE_DIR=1 $SDK/.venv/bin/python -m pip install \\\n  --index-url https://stable.repo.amd.com/rocm/whl-next/ \"rocm[libraries,devel,device-gfx1151]==10.0.0\"\n$SDK/.venv/bin/rocm-sdk init\n\nR=$($SDK/.venv/bin/rocm-sdk path --root)\nexport PATH=\"$R/bin:$PATH\" HIP_PATH=\"$R\" ROCM_PATH=\"$R\"\ncd $LLMS/gufo\ncmake --preset release -B build/release-rocm10 \\\n  -DCMAKE_PREFIX_PATH=\"$R\" -DCMAKE_HIP_COMPILER=\"$R/lib/llvm/bin/clang++\" \\\n  \"-DCMAKE_BUILD_RPATH=$R/lib;$R/lib/rocm_sysdeps/lib;$R/lib/llvm/lib\"\ncmake --build build/release-rocm10 -j 16\n./build/release-rocm10/gufo diagnose\n```\n\n- **GCC 16 or newer** (Arch, for example) breaks ROCm clang HIP compiles in`<format>` . Install GCC 15 and add the following to the`cmake --preset` line:\n\n```\n-DCMAKE_C_COMPILER=gcc-15 -DCMAKE_CXX_COMPILER=g++-15 -DCMAKE_HIP_FLAGS=--gcc-install-dir=$(dirname $(gcc-15 -print-libgcc-file-name))\n```\n\n- The RPATH points the binary at the SDK's libraries, so it runs without any env vars.\n- **Fallback:** gufo's own qualified toolchain is a system ROCm 7.2.3 with GCC 15.3. Follow \"Build from source / Without Nix\" in gufo's README. On the reference machine, ROCm 7.2.4 and ROCm 10 builds were within 3% of each other.\n\n1. Stop the daemon first: `llamastash daemon stop` . The config is read only at start.\n2. Merge the YAML below into `~/.config/llamastash/config.yaml` , keeping anything already there (for example`backend.llamacpp` ).\n3. Replace every `/ABS/...` path with a real absolute path:\n  - `/ABS/llms` =`$LLMS` from step 5\n  - `/ABS/hub` = the`hub/` directory from step 3\n4. Start the daemon again: `llamastash daemon start` .\n\nPresets set only the context window, plus the KV pool for Halogen. Thinking effort is not a preset. Each engine defaults to `xhigh`, and a client that sends `reasoning_effort` (such as Pi's thinking level or Claude Code's `/effort`) overrides it per request.\n\n```\nproxy:\n  idle_ttl_secs: 0 # a loaded model stays until you stop it\nbackend:\n  generic:\n    servers:\n      # Server `generic-gufo` on the Flash-Next GGUF row. --reasoning-effort is\n      # the server default; a request's reasoning_effort overrides it.\n      - name: gufo\n        model: Qwen3.8-Flash-Next-UD-Q4_K_XL\n        binary: /ABS/llms/gufo/build/release-rocm10/gufo\n        args:\n          [\n            serve,\n            --host,\n            \"{host}\",\n            --port,\n            \"{port}\",\n            --sessions,\n            \"1\",\n            llm,\n            --served-model-name,\n            \"{name}\",\n            --model,\n            \"{model}\",\n            --mtp-model,\n            /ABS/hub/models--unsloth--Qwen3.8-Flash-Next-GGUF/snapshots/38bb39ee97821de2c9009abb7e93950eec396e66/MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf,\n            --top-p,\n            \"0.95\",\n            --top-k,\n            \"20\",\n            --reasoning-effort,\n            xhigh,\n          ]\n        knobs:\n          - { flag: --context, ctx: true, default: \"131072\" }\n          - { flag: --speculative, default: mtp }\n          - { flag: --temperature, default: \"1.0\" }\n          - --seed\n        ready: /ready\n        stop_grace_secs: 60\n        # gufo 404s on any model name but --served-model-name.\n        rewrite_model: true\n      # Server `generic-gufo-27b` on the 27B UD-Q6_K row.\n      - name: gufo-27b\n        model: Qwen3.8-27B-UD-Q6_K\n        binary: /ABS/llms/gufo/build/release-rocm10/gufo\n        args:\n          [\n            serve,\n            --host,\n            \"{host}\",\n            --port,\n            \"{port}\",\n            --sessions,\n            \"1\",\n            llm,\n            --served-model-name,\n            \"{name}\",\n            --model,\n            \"{model}\",\n            --speculative,\n            dflash2,\n            --dflash-model,\n            /ABS/hub/models--z-lab--Qwen3.8-27B-DFlash2-GGUF/snapshots/2d9571f8ce46e151f61c6499c99dee6079e1d610/Qwen3.8-27B-DFlash2-Q4_K_M.gguf,\n            --top-p,\n            \"0.95\",\n            --top-k,\n            \"20\",\n            --reasoning-effort,\n            xhigh,\n          ]\n        knobs:\n          - { flag: --context, ctx: true, default: \"131072\" }\n          - { flag: --temperature, default: \"1.0\" }\n          - --seed\n        ready: /ready\n        stop_grace_secs: 60\n        rewrite_model: true\n      # Own row (weights are not a GGUF). Docker wrapper from step 4.\n      - name: flash-next-halogen\n        binary: ~/.local/bin/halogen-serve\n        args: [\"{port}\"]\n        arch: qwen4exp\n        params: 177B\n        quant: v2-4.16bpw\n        ctx: 262144\n        # No GGUF to read these from, so the row declares them for `integrations`.\n        # Thinking off plus the template's levels; leave `vision` off: the wrapper\n        # doesn't set HALOGEN_VISION_TOWER, so Halogen refuses images.\n        reasoning_effort: [none, low, medium, xhigh]\n        reasoning_effort_default: xhigh # same as halogen-effort's default\n        # At the 262144 window and pool Halogen holds about 87 GiB. Also gates\n        # admission. A 786432 pool adds ~14.4 GiB this does not count.\n        memory_gib: 87\n        knobs:\n          - {\n              flag: --ctx-window,\n              id: halogen-ctx,\n              ctx: true,\n              default: \"131072\",\n            }\n          # Positions resident across all sessions, reserved at start (~7.2 GiB\n          # per 262144) out of the lookup table's page cache. Must be >= ctx.\n          - { flag: --halogen-kv-pool, id: halogen-pool, default: \"131072\" }\n          - { flag: --halogen-temperature, id: halogen-temp, default: \"1.0\" }\n          - { flag: --halogen-mtp-depth, id: halogen-mtp-depth, default: \"3\" }\n          # Server default only; a request's reasoning_effort overrides it.\n          - {\n              flag: --halogen-reasoning-effort,\n              id: halogen-effort,\n              default: xhigh,\n            }\n        env:\n          HALOGEN_MODEL_ID: \"{name}\"\n          HALOGEN_CTX: \"{halogen-ctx}\"\n          HALOGEN_KV_POOL_POSITIONS: \"{halogen-pool}\"\n          HALOGEN_TEMPERATURE: \"{halogen-temp}\"\n          HALOGEN_MTP_DEPTH: \"{halogen-mtp-depth}\"\n          HALOGEN_REASONING_EFFORT: \"{halogen-effort}\"\n        ready: /v1/models\n        stop_grace_secs: 90\n        ready_timeout_secs: 600\npresets:\n  # 27B on gufo with the DFlash2 draft, model-card sampling. Give clients a\n  # max_tokens of 16k or more, or long thinking ends in an empty answer.\n  Qwen3.8-27B-UD-Q6_K.gguf:\n    default: 256k\n    entries:\n      128k:\n        server: generic-gufo-27b\n        knobs: { context: 131072, temperature: \"1.0\" }\n      256k:\n        server: generic-gufo-27b\n        knobs: { context: 262144, temperature: \"1.0\" }\n  # Flash-Next GGUF on gufo with MTP. Stock llama.cpp can't load the qwen4exp MTP head.\n  Qwen3.8-Flash-Next-*:\n    default: 128k\n    entries:\n      128k:\n        server: generic-gufo\n        knobs: { context: 131072, speculative: mtp, temperature: \"1.0\" }\n      256k:\n        server: generic-gufo\n        knobs: { context: 262144, speculative: mtp, temperature: \"1.0\" }\n  # Halogen native v2, MTP depth 3, model-card sampling: the fastest Flash-Next\n  # setup. The pool decides how many sessions stay resident; one that doesn't\n  # fit re-reads its whole prompt every turn.\n  flash-next-halogen:\n    default: solo-256k\n    entries:\n      solo-256k:\n        knobs: { halogen-ctx: 262144, halogen-pool: 262144 }\n      dual-128k:\n        knobs: { halogen-ctx: 131072, halogen-pool: 262144 }\n      triple-128k:\n        knobs: { halogen-ctx: 131072, halogen-pool: 393216 }\n      dual-256k:\n        knobs: { halogen-ctx: 262144, halogen-pool: 524288 }\n      # Three full-length sessions; ~14.4 GiB less page cache than solo-256k.\n      triple-256k:\n        knobs: { halogen-ctx: 262144, halogen-pool: 786432 }\n```\n\n1. Run `llamastash list` . You should see:\n  - the two GGUF rows with backend `llamacpp|generic`\n  - a `flash-next-halogen` row with ctx 262144 and size 87G\n  - `llamastash presets list <model>` showing the presets above\n2. the two GGUF rows with backend \n3. Test each engine one at a time. Start it, send one chat request through the proxy, then stop it:\n  - `llamastash start flash-next-halogen --preset solo-256k`\n  - `llamastash start Qwen3.8-Flash-Next-UD-Q4_K_XL --preset 128k`\n  - `llamastash start Qwen3.8-27B-UD-Q6_K --preset 128k`\n  - The proxy is at `http://$(llamastash status --json | jq -r .proxy.listen)/v1` . The model id is`<id>@<preset>` , as listed by`GET /v1/models` .\n  - Stop with `llamastash stop <model>` . Wait for the stop to finish before starting the next engine.\n4. Watch `llamastash logs <model>` . Expected:\n  - Halogen cold load takes 85 to 105 s, or about 6 s when the weights are still in page cache. gufo takes about 22 s.\n  - Halogen's startup log prints a memory line. If it warns about **compaction stalls** , run`echo 1 | sudo tee /proc/sys/vm/compact_memory` before the next start. Its log suggests this instead of a reboot.\n  - After any unclean exit, check `cat /sys/class/drm/card*/device/mem_info_gtt_used` . Tens of GiB used with nothing running means the GPU kept the memory, and only a reboot frees it.\n5. Reference numbers from an ASUS ROG Flow Z13 (Strix Halo 128 GB) at a 70 W TDP, 64k window with a 32k prompt, temperature 1.0 / top_p 0.95 / top_k 20. A desktop at 85 to 100 W should be faster.\n\n| Engine | Prefill t/s | Decode t/s | Draft accept | \n|---|---|---|---|\n| Halogen 0.15.1, v2 | 1,191 | 39.4 | 85% | \n| gufo 0.3.0, Flash-Next, ROCm 10 | 1,042 | 34.8 | 75% | \n| gufo, 27B UD-Q6_K + DFlash2 | 405 | 17 to 21 | 46% | \n\nHalogen 0.16.1 on the same laptop (performance profile, on AC), with a 30k-token wikitext prompt at the 256k window: prefill 1,276 to 1,362 t/s, decode 42.3 to 43.4 t/s, draft accept 78 to 81%. The prompt differs from the table's, so this is not a direct comparison.\n\nLlamaStash's `integrations` command registers every **favorite** model, one entry per preset, as `<id>@<preset>`. So mark the three rows as favorites first:\n\n```\nllamastash favorites add Qwen3.8-Flash-Next-UD-Q4_K_XL\nllamastash favorites add Qwen3.8-27B-UD-Q6_K\nllamastash favorites add flash-next-halogen\nllamastash integrations pi opencode claude-code\n```\n\n- **Pi:** patches`~/.pi/agent/models.json` with a`llamastash` provider. The API key resolves through`!llamastash api-key` . Pick a model like`flash-next-halogen@solo-256k` in Pi.\n- **OpenCode:** patches its config with the proxy URL and the same model list.\n- **Claude Code:** writes`~/.config/llamastash/claude-code.sh` and doesn't touch`~/.claude/settings.json` . Run it with`source ~/.config/llamastash/claude-code.sh && claude` .\n  - Use **Halogen** for Claude Code. Set both`ANTHROPIC_MODEL` and`ANTHROPIC_SMALL_FAST_MODEL` in that file to a Halogen preset id, such as`flash-next-halogen@solo-256k` .\n  - gufo can't serve Claude Code: as of v0.5.0 its `/v1/messages` rejects tools,`thinking` ,`output_config` and streaming.\n- Use \n- Add `codex` ,`aider` ,`continue` ,`zed` or`env-sh` to the same command for other tools.\n- Re-run `llamastash integrations ...` after you add favorites or presets.\n- **Effort and image fields.**`integrations` writes them for every row:`reasoning` , plus`thinkingLevelMap` in Pi or`variants` in OpenCode, with levels off (`none` ), low, medium and xhigh.\n  - The gufo rows get theirs from the GGUF chat template, and image input from the `mmproj` file in step 3.\n  - The Halogen row gets effort from the `reasoning_effort` fields in step 6 and no image input.\n  - To check, open Pi's `/model` and pick a`flash-next-halogen@...` entry. Its thinking-level control should list off, low, medium and xhigh.\n  - gufo needs `mmproj-BF16.gguf` for images. The integration only checks that some projector file is present, so a row can list image input while gufo still refuses images.\n- The gufo rows get theirs from the GGUF chat template, and image input from the \n\nTell the user:\n\n- the versions installed\n- anything you changed from this guide, and why\n- any host check that failed\n- the prefill and decode numbers from step 7", "url": "https://wpnews.pro/news/handover-halogen-gufo-behind-llamastash-on-strix-halo-qwen3-8-flash-next-and-27b", "canonical_source": "https://gist.github.com/deepu105/d0b321f3256edf67ebd85e747c0010b9", "published_at": "2026-10-02 13:33:54+00:00", "updated_at": "2026-10-04 20:12:28.372566+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-tools", "mlops", "developer-tools"], "entities": ["LlamaStash", "Halogen", "gufo", "Qwen3.8 Flash-Next", "Unsloth", "AMD Strix Halo", "Ryzen AI Max+ 395", "Claude Code"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/handover-halogen-gufo-behind-llamastash-on-strix-halo-qwen3-8-flash-next-and-27b", "markdown": "https://wpnews.pro/news/handover-halogen-gufo-behind-llamastash-on-strix-halo-qwen3-8-flash-next-and-27b.md", "text": "https://wpnews.pro/news/handover-halogen-gufo-behind-llamastash-on-strix-halo-qwen3-8-flash-next-and-27b.txt", "jsonld": "https://wpnews.pro/news/handover-halogen-gufo-behind-llamastash-on-strix-halo-qwen3-8-flash-next-and-27b.jsonld"}}