{"slug": "ryzen-ai-halo-halogen-server-testing-notes", "title": "Ryzen AI Halo: Halogen Server Testing Notes", "summary": "A community tester published deployment notes for the Halogen Flash Server, an optimized server for running Qwen3.8-Flash-Next on AMD Strix Halo (gfx1151) hardware, reporting a roughly 10% performance uplift from the amd_iommu=off kernel parameter. The notes cover a real-world deployment on an AMD Ryzen AI MAX+ 395 mini-PC with 128 GB unified memory, using Debian 13, Podman, and ROCm 7.14, and walk through stock launch, power tuning, YaRN 1M context extension, and container management. The model weights and n-gram table total about 111 GB, split between a 63 GiB main checkpoint, a 48 GiB n-gram lookup table for speculative decoding, and an optional 857 MiB vision sidecar.", "body_md": "# \n\nI dunno that this will become a video, theres *so many* videos but I wanted to share this with the community as its my notes from experimenting with the “Halogen” engine on the Strix Halo platform.\n\nMaybe I’ll tweet about it.\n\nWhat’s Halogen? It’s this sort of interesting optimized server for running qwen 3.8 flash next on Strix halo. You should read about the project, it’s interesting.\n\nThis was tested on several Strix Halo platforms including the ASUS ROG Flow Z13.\n\n**Original project**: [github.com/peonist-ai/halogen-flash-server](https://github.com/peonist-ai/halogen-flash-server) — the fastest way to run Qwen3.8-Flash-Next on AMD Strix Halo (gfx1151).\n\n**Model**: [peonist-ai/halogen-qwen3.8-flash-next](https://huggingface.co/peonist-ai/halogen-qwen3.8-flash-next) on Hugging Face.\n\n**Platform**: AMD Ryzen AI MAX+ 395 (Radeon 8060S, 128 GB unified memory).\n\nThis guide covers a real-world deployment on a Strix Halo mini-PC. Normally I use Ubuntu LTS with Docker, but I decided to try Podman on Debian 13 this time around.\n\nWe walk through the stock launch, power tuning, YaRN 1M context extension, and the `amd_iommu=off` kernel parameter that’s good for about +10% perf uplift. Each step includes the commands, the reasoning, and the performance impact.\n\n## \n\n1. [Prerequisites & Hardware](#1-prerequisites--hardware)\n2. [Stock Launch (Reference)](#2-stock-launch-reference)\n3. [Power Tuning: Platform Profile & CPU Governor](#3-power-tuning--platform-profile--cpu-governor)\n4. [YaRN 1M Context Extension](#4-yarn-1m-context-extension)\n5. [amd_iommu=off – The Secret Sauce](#5-amd_iommuoff--the-secret-sauce)\n6. [Container Management](#6-container-management)\n7. [Performance Results](#7-performance-results)\n8. [Troubleshooting](#8-troubleshooting)\n9. [Appendix: Rollback & Recovery](#9-appendix-rollback--recovery)\n\n## \n\n### \n\nAMD Ryzen AI Halo\n\nGMKtec\n\nROG Flow Z13\n\nFramework 13\n\nThese systems are all bananas. The software has made the story. This should all work on AMD’s **Gorgon Halo** as well – 400 series with 192gb ram.\n\nThese all had 128gb LPDDR5 unified memory.\n\n| Component | Detail | \n| **CPU/APU** | AMD Ryzen AI MAX+ 395 (16C/32T) | \n| **GPU** | Radeon 8060S (gfx1151, integrated, 128 GB unified memory) | \n| **RAM** | 128 GB LPDDR5X (shared CPU/GPU pool) | \n| **Storage** | NVMe SSD | \n| **Platform** | Mini-PC (ASUS ROG Flow Z13) | \n\n ### \n\n- **OS** : Debian 13 “Trixie” (or any recent Linux with ROCm 7.14 support)\n- **Container runtime** : Podman (Docker works too; adjust`--group-add` accordingly)\n- **ROCm** : 7.14 (shipped with the halogen-flash-server image)\n- **Kernel** : 6.18.44+rex+5-amd64 (vendor kernel from AMD on the AI Halo, but CachyOS kernel tested and works too.)\n\n### \n\nDownload the model weights and n-gram table from Hugging Face (~111 GB total):\n\n```\n# Install huggingface-cli if needed\npip install huggingface-hub\n\n# Download model\nhuggingface-cli download peonist-ai/halogen-qwen3.8-flash-next \\\n  --local-dir ~/halogen-models\n```\n\n^ strictly speaking this isn’t needed as the halogen github link, at first run, downloads the model. I like to do this to ensure that “extra” downloads across test runs do not happen. Save my precious bandwidth.\n\nExpected files:\n\n| File | Size | Purpose | \n| `qwen38-flash-next-v2.hgn` | 63 GiB | Main checkpoint (weights) | \n| `qwen38-flash-next-ngram.hgn` | 48 GiB | N-gram lookup table for speculative decoding | \n| `qwen38-flash-next-vision.hgn` | 857 MiB | Vision sidecar (optional, not loaded by default) | \n| `tokenizer/` | — | Tokenizer files | \n\n \n\nIf you aren’t read in on the whole N-gram thing, this is a “new” thing where part of the model is designed to run from dram at dram speeds. This is a unified memory platform so it manifests as a speed benefit, but the N-gram model work should be exciting to everyone since it promises decent performance on “normal” systems that have fast vram (faster than LPDDR5) and slower system memory.\n\n## \n\nThis is the baseline from the [GitHub README](https://github.com/peonist-ai/halogen-flash-server). Run the container as-is, no tuning:\n\n```\npodman run -d --name halogen-flash-server \\\n  -p 8731:8731 \\\n  --device /dev/kfd --device /dev/dri \\\n  --group-add keep-groups \\\n  --ipc=host \\\n  --ulimit memlock=-1:-1 \\\n  -v ~/halogen-models:/models \\\n  ghcr.io/peonist-ai/halogen-flash-server:0.15.0\n```\n\nThe server listens on port 8731. Check health:\n\n```\ncurl -s http://localhost:8731/health | python3 -m json.tool\n```\n\n**Key observations at stock:**\n\n- Platform profile: `balanced`\n- CPU governor: `powersave`\n- GPU clocks: ~2000 MHz under load\n- GPU power: ~60W\n- Context: 262,144 tokens (native, no YaRN)\n- Output cap: 65,536 tokens\n\n## \n\nThe stock `balanced` profile and `powersave` governor leave a lot of performance on the table. I suspect some of the benchmarks posted on github on strix halo were run with a conservative CPU governor. The GPU can hit 2900 MHz and 130W, but the power envelope needs to be opened up. The GMKtek has been tuned for > 150W, too, but I have omitted the results from this writeup to keep it from being confusing (it was only worth about 2% more, at best, performance).\n\n### \n\n```\ncat /sys/firmware/acpi/platform_profile\n# → balanced\n\ncat /sys/devices/system/cpu/cpufreq/policy*/scaling_governor | sort -u\n# → powersave\n\ncat /sys/class/drm/card0/device/pp_dpm_sclk\n# 0: 600Mhz *\n# 1: 1100Mhz\n# 2: 2900Mhz\n```\n\n### \n\n```\n# Set platform profile to performance\nsudo -S sh -c 'echo performance > /sys/firmware/acpi/platform_profile'\n\n# Set all CPU governors to performance\nfor f in /sys/devices/system/cpu/cpufreq/policy*/scaling_governor; do\n  sudo -S sh -c \"echo performance > \\\"$f\\\"\"\ndone\n```\n\n### \n\n```\ncat /sys/firmware/acpi/platform_profile    # → performance\ncat /sys/devices/system/cpu/cpufreq/policy0/scaling_governor  # → performance\n```\n\n### \n\n| Parameter | Before | After | \n| Platform profile | `balanced` | `performance` | \n| CPU governor | `powersave` | `performance` | \n| GPU clock under load | ~2000 MHz | ~2850 MHz | \n| GPU power under load | ~60W | ~130W | \n\n ### \n\n```\n   sudo -S sh -c 'echo balanced > /sys/firmware/acpi/platform_profile'\nfor f in /sys/devices/system/cpu/cpufreq/policy*/scaling_governor; do\n   sudo -S sh -c \"echo powersave > \\\"$f\\\"\"\ndone\n```\n\n**Note**: These settings are ephemeral — they reset on reboot. See the [Post-Reboot Automation](#post-reboot-automation) section to persist them.\n\n## \n\nThe model’s native context is 262,144 tokens. With YaRN (Yet another RoPE extensioN) at factor 4, we extend this to 1,048,576 tokens (~1M). This requires a larger KV cache pool — about 28.8 GiB of the 128 GiB unified memory.\n\n### \n\n```\npodman run -d --name halogen-flash-server \\\n  -p 8731:8731 \\\n  --device /dev/kfd --device /dev/dri \\\n  --group-add keep-groups \\\n  --ipc=host \\\n  --ulimit memlock=-1:-1 \\\n  -e HALOGEN_ROPE_YARN=4 \\\n  -e HALOGEN_CTX=1048576 \\\n  -e HALOGEN_MAX_THINKING_TOKENS=262144 \\\n  -e HALOGEN_MAX_TOKENS_DEFAULT=393216 \\\n  -e HALOGEN_MAX_TOKENS_CAP=393216 \\\n  -v ~/halogen-models:/models \\\n  ghcr.io/peonist-ai/halogen-flash-server:0.15.0\n```\n\n### \n\n| Variable | Value | Purpose | \n| `HALOGEN_ROPE_YARN=4` | 4 | YaRN scaling factor. 4 × 262K native = 1,048,576 context | \n| `HALOGEN_CTX=1048576` | 1,048,576 | Max context length (KV pool positions) | \n| `HALOGEN_MAX_THINKING_TOKENS=262144` | 262,144 | Reasoning budget cap (256K tokens for thinking) | \n| `HALOGEN_MAX_TOKENS_DEFAULT=393216` | 393,216 | Default total output budget (thinking + answer) | \n| `HALOGEN_MAX_TOKENS_CAP=393216` | 393,216 | Hard cap; requests above get 400 error | \n\n ### \n\n``` python\ncurl -s http://localhost:8731/health | python3 -c \"\nimport sys, json\nd = json.load(sys.stdin)\nprint('Context:', d['context'])\nprint('YaRN:', d.get('rope_scaling'))\nprint('Max tokens default:', d.get('max_tokens_default'))\nprint('Max thinking tokens:', d.get('max_thinking_tokens_default'))\n\"\n```\n\nExpected output:\n\n```\nContext: 1048576\nYaRN: {'type': 'yarn', 'factor': 4.0, 'original_context': 262144}\nMax tokens default: 393216\nMax thinking tokens: 262144\n```\n\n### \n\n```\nrope: static YaRN, factor 4 over the native 262144 (attention scale 1.138629)\nstartup [   3.6 s] KV pool reserved: 1048576 positions (about 28.8 GiB)\nstartup [   3.7 s] memory: 62.1 GiB of weights locked in RAM, 28.8 GiB of KV pool, 8.1 GiB of working memory, 99.0 GiB in all\nstartup [   3.7 s] host memory left for everything else: ~20 GiB\n```\n\n### \n\nWith YaRN 1M enabled, the server uses ~99 GiB of the 128 GiB pool:\n\n- **62 GiB** — model weights (pinned in RAM)\n- **29 GiB** — KV cache (1,048,576 positions)\n- **8 GiB** — working memory\n- **~20 GiB** — remaining for the OS and other services\n\nIf you’re tight on memory, you can halve the KV pool:\n\n```\n-e HALOGEN_KV_POOL_POSITIONS=524288\n-e HALOGEN_CTX=524288\n```\n\nThis reduces context to 524K but frees ~14 GiB.\n\n## \n\nThe single biggest performance unlock on Strix Halo is disabling the IOMMU. On this platform, the IOMMU adds overhead to GPU memory access that visibly impacts prefill throughput. The project’s [reference machine](https://github.com/peonist-ai/halogen-flash-server) runs with `amd_iommu=off` as part of its kernel command line, and the README notes it’s worth **13–16% of prefill performance**.\n\n**Note**: The reference machine also uses additional GPU-tuning kernel parameters: `amdgpu.vm_update_mode=0 amdgpu.noretry=0 amdgpu.gttsize=126976 ttm.pages_limit=32505856 amdgpu.sg_display=0`. We did NOT apply these in our testing — our results come from `amd_iommu=off` alone, combined with the userspace power tuning from Section 3.\n\n### \n\nThe system uses **systemd-boot** (not GRUB). Boot entries live on the EFI System Partition at `/efi/loader/entries/`.\n\n**Step 1: Identify the current boot entry**\n\n```\nbootctl status | grep \"Current Entry\"\n# → amd-ryzen-ai-developer-platform-6.18.44+rex+5-amd64.conf\n```\n\n**Step 2: Back up the entry**\n\n```\nsudo cp /efi/loader/entries/amd-ryzen-ai-developer-platform-6.18.44+rex+5-amd64.conf \\\n       /efi/loader/entries/amd-ryzen-ai-developer-platform-6.18.44+rex+5-amd64.conf.bak-iommu\n```\n\n**Step 3: Add `amd_iommu=off` to the kernel command line**\n\n```\nsudo sed -i 's/^options /options amd_iommu=off /' \\\n  /efi/loader/entries/amd-ryzen-ai-developer-platform-6.18.44+rex+5-amd64.conf\n```\n\n**Step 4: Verify the edit**\n\n```\ncat /efi/loader/entries/amd-ryzen-ai-developer-platform-6.18.44+rex+5-amd64.conf\n```\n\nShould show:\n\n```\noptions amd_iommu=off     root=UUID=... splash quiet loglevel=3\n```\n\n**Step 5: Reboot**\n\n```\nsudo systemctl reboot\n```\n\n**Step 6: Verify after reboot**\n\n```\ncat /proc/cmdline | tr ' ' '\\n' | grep iommu\n# → amd_iommu=off\n\nsudo dmesg | grep -c AMD-Vi\n# → 0  (IOMMU is disabled)\n\nsudo dmesg | grep -i kfd\n# → kfd kfd: amdgpu: added device 1002:1586  (KFD still works)\n```\n\n### \n\nYes. On this kernel (6.18.44+rex+5-amd64), the AMD KFD (Kernel Fusion Driver) falls back to GART-based initialization when the IOMMU is disabled. You’ll see this in dmesg:\n\n```\nkfd kfd: amdgpu: Allocated 3969056 bytes on gart\nkfd kfd: amdgpu: Total number of KFD nodes to be created: 1\nkfd kfd: amdgpu: added device 1002:1586\n```\n\nROCm 7.14 works normally. The container needs `--group-add` with the render group’s GID (typically 992) because rootless Podman remaps device ownership, but that’s a container-runtime detail, not an IOMMU issue.\n\n### \n\nIf the system fails to boot after adding `amd_iommu=off`, use the boot menu to select the backup entry or use a recovery USB to restore the backup:\n\n```\n# From a recovery shell:\nsudo cp /efi/loader/entries/...conf.bak-iommu /efi/loader/entries/...conf\n```\n\n## \n\n### \n\nOn this Debian 13 system, **rootless Podman remaps device node ownership** inside the container. The `/dev/kfd` and `/dev/dri/renderD128` devices show up as owned by `nobody:nogroup`, which breaks ROCm’s permission check.\n\n**Fix**: Run the container under `sudo podman` with the render group’s GID:\n\n```\n# Find the render GID\ngrep ^render /etc/group | cut -d: -f3\n# → 992\n\n# Launch with sudo\nsudo podman run -d --name halogen-flash-server \\\n  -p 8731:8731 \\\n  --device /dev/kfd --device /dev/dri \\\n  --group-add 992 \\\n  --ipc=host \\\n  --ulimit memlock=-1:-1 \\\n  -e HALOGEN_ROPE_YARN=4 \\\n  -e HALOGEN_CTX=1048576 \\\n  -e HALOGEN_MAX_THINKING_TOKENS=262144 \\\n  -e HALOGEN_MAX_TOKENS_DEFAULT=393216 \\\n  -e HALOGEN_MAX_TOKENS_CAP=393216 \\\n  -v /home/w/halogen-models:/models \\\n  ghcr.io/peonist-ai/halogen-flash-server:0.15.0\n```\n\n### \n\nSave this as `~/halogen-podman` and `chmod +x`:\n\n``` bash\n#!/bin/bash\nPASS=your_sudo_password_here\nACTION=\"$1\"\nshift\ncase \"$ACTION\" in\n  start|stop|restart|logs|rm)\n    echo \"$PASS\" | sudo -S podman \"$ACTION\" halogen-flash-server \"$@\"\n    ;;\n  ps)\n    echo \"$PASS\" | sudo -S podman ps -a --filter name=halogen-flash-server \"$@\"\n    ;;\n  health)\n    echo \"$PASS\" | sudo -S curl -s http://localhost:8731/health \"$@\"\n    ;;\n  run)\n    echo \"$PASS\" | sudo -S podman run -d --name halogen-flash-server \\\n      -p 8731:8731 \\\n      --device /dev/kfd --device /dev/dri \\\n      --group-add 992 \\\n      --ipc=host \\\n      --ulimit memlock=-1:-1 \\\n      -e HALOGEN_ROPE_YARN=4 \\\n      -e HALOGEN_CTX=1048576 \\\n      -e HALOGEN_MAX_THINKING_TOKENS=262144 \\\n      -e HALOGEN_MAX_TOKENS_DEFAULT=393216 \\\n      -e HALOGEN_MAX_TOKENS_CAP=393216 \\\n      -v /home/w/halogen-models:/models \\\n      ghcr.io/peonist-ai/halogen-flash-server:0.15.0\n    ;;\n  *)\n    echo \"Usage: halogen-podman {start|stop|restart|logs|rm|ps|health|run}\"\n    ;;\nesac\n```\n\nUsage:\n\n```\n./halogen-podman start    # Start the container\n./halogen-podman stop     # Stop it\n./halogen-podman restart  # Restart it\n./halogen-podman logs     # View logs\n./halogen-podman health   # Check health endpoint\n./halogen-podman ps       # Container status\n./halogen-podman run      # Create a fresh container\n```\n\nThere is probably a better way to do this. This kind of thing isn’t necessary with Docker, or I might need more XP with podman to understand what I’ve missed. I suspect that because when you’re in the docker group, you’ve basically got root, and this is a different architectural choice in podman to prevent that. But I need the hardware. So this is just making some notes for myself about that.\n\n### \n\nAfter every reboot, you need to:\n\n1. Re-apply the performance profile + governor\n2. Start the halogen container\n\nA simple systemd oneshot service can handle this kind of thing, though. Create `/etc/systemd/system/halogen-tune.service`:\n\n```\n[Unit]\nDescription=Halogen power tuning\nAfter=multi-user.target\n\n[Service]\nType=oneshot\nExecStart=/usr/local/bin/halogen-tune.sh\nRemainAfterExit=yes\n\n[Install]\nWantedBy=multi-user.target\n```\n\nAnd `/usr/local/bin/halogen-tune.sh`:\n\n``` bash\n#!/bin/bash\necho performance > /sys/firmware/acpi/platform_profile\nfor f in /sys/devices/system/cpu/cpufreq/policy*/scaling_governor; do\n  echo performance > \"$f\"\ndone\nsudo systemctl enable halogen-tune.service\n```\n\nThen add the container start to your crontab (`@reboot /home/w/halogen-podman start`) or another systemd unit.\n\nDifferent distros handle this in different ways. I’m getting a bit rusty on Debian and there may be a less brute-force way to do this. But this is good practice for you sysadmin-in-training folks out there.\n\n## \n\n### \n\nWe used the built-in `sweep` command, which measures end-to-end HTTP request latency including prefill and decode:\n\n```\n# Prefill benchmarks (3 repetitions)\nsudo podman run --rm ... halogen-flash-server:0.15.0 sweep -p 8192,32768 -n 128 -r 3\nsudo podman run --rm ... halogen-flash-server:0.15.0 sweep -p 131072 -n 128 -r 1\n\n# Decode benchmark (10 prompt shapes, 1 repetition)\nsudo podman run --rm ... halogen-flash-server:0.15.0 bench mtp 256 low 1\n```\n\nAll benchmarks use the MTP (Multi-Token Prediction) drafter, which is the default and fastest mode.\n\n### \n\n| Test | Ref (0.14.1) | Baseline (0.14.0) | Pass 1 (stock) | Pass 2 (tuned) | Pass 3 (tuned+IOMMU) | \n| **pp8192** | 1,584 t/s | 1,268 t/s | 1,350 t/s | 1,679 t/s | **1,768 t/s** | \n| **pp32768** | 1,567 t/s | 1,442 t/s | 1,216 t/s | 1,469 t/s | **1,695 t/s** | \n| **pp131072** | 1,517 t/s | 1,390 t/s | 1,171 t/s | 1,426 t/s | **1,621 t/s** | \n| **tg128 MTP** | 46.0 t/s | 39.9 t/s | 48.1 t/s | 52.0 t/s | 52.2 t/s | \n| **mtp 256 low 1** | — | — | — | — | **53.4 t/s** | \n\n *pp = prefill (prompt processing), tg = token generation. All values in tokens/second. Higher is better. Reference and Baseline columns from the [project README](https://github.com/peonist-ai/halogen-flash-server).*\n\n### \n\n| Pass | Changes | \n| **Ref (0.14.1)** | Project’s reference machine – ROCm 7.14, IOMMU=off, additional GPU kernel params, ~85W sustained | \n| **Baseline (0.14.0)** | Same machine, prior software version (for comparison) | \n| **Pass 1** | Our stock baseline – balanced profile, powersave governor, IOMMU on | \n| **Pass 2** | Platform profile → `performance` , CPU governor →`performance` on all 32 cores | \n| **Pass 3** | Pass 2 + `amd_iommu=off` kernel parameter + reboot | \n\n ### \n\nMost real-world usage will be at 256K context or less – the model’s native length. The YaRN 1M config is available for deep-context tasks, but 256K covers the vast majority of chat, coding, and analysis workloads. Here’s how performance compares:\n\n| Metric | 256K Context (native) | 1M Context (YaRN 4x) | \n| KV pool size | ~7 GiB | ~29 GiB | \n| Total memory used | ~77 GiB | ~99 GiB | \n| Host memory free | ~42 GiB | ~20 GiB | \n| Prefill @ 8K | **~1,770 t/s** | ~1,770 t/s | \n| Prefill @ 32K | **~1,700 t/s** | ~1,700 t/s | \n| Prefill @ 131K | **~1,620 t/s** | ~1,620 t/s | \n| Decode (MTP, short ctx) | **~53 t/s** | ~53 t/s | \n| Decode (MTP, deep ctx ~260K) | — | ~45 t/s | \n\n The prefill and short-context decode numbers are nearly identical between configs – the YaRN scaling doesn’t add meaningful overhead for prompt processing or short generations. The difference appears at deep context: the README reports 45.0 tok/s decode at 258K context with the 1M config, vs 46.0 tok/s at 32K context. The 1M config also enables prompt caching across very long sessions (the follow-up turn at 100K context is ~2s).\n\n**Recommendation**: Run with the 1M config by default. The memory cost (~22 GiB extra for the larger KV pool) is worth the flexibility, and performance at 256K seems basically identical. Only drop to 524K or 256K if you’re running other memory-hungry services alongside the server.\n\n### \n\n| Metric | Stock (Pass 1) | Tuned (Pass 2) | Tuned+IOMMU (Pass 3) | Ref Machine | \n| GPU clock under load | ~2000 MHz | ~2850 MHz | ~2900 MHz | — | \n| GPU power under load | ~60W | ~130W | ~140W | ~85W | \n| Idle clock | 600 MHz | 600 MHz | 600 MHz | — | \n| Idle power | ~5W | ~5W | ~5W | — | \n\n The GPU hits its 2900 MHz ceiling and 140W power limit (160w on the GMKtek) after tuning. The reference machine runs at ~85W sustained – our higher power draw is expected given the `performance` profile (the reference likely uses a tuned `balanced` profile with IOMMU=off and the additional GPU kernel parameters). The remaining small gap vs reference at 131K prefill is likely memory-bandwidth bound rather than clock-limited.\n\n### \n\n```\npp8192:     ████████████████████░░░░░░░░░░  1,768 t/s  (+12% vs ref 1,584)\npp32768:    █████████████████████░░░░░░░░░░  1,695 t/s  (+8% vs ref 1,567)\npp131072:   ████████████████████████░░░░░░░░  1,621 t/s  (+7% vs ref 1,517)\nDecode:     ██████████████████████████░░░░░░  53.4 t/s  (+16% vs ref 46.0)\n```\n\n### \n\n- **Pass 3 (tuned + IOMMU=off) exceeds the reference** at every prefill size, despite not using the additional GPU kernel parameters (`amdgpu.vm_update_mode=0` , etc.) that the reference machine employs. This suggests the userspace power tuning (performance profile + governor) is doing significant work beyond just IOMMU=off.\n- **Pass 1 was below the 0.14.0 baseline** — our stock configuration (IOMMU on, balanced profile) was leaving ~20% on the table.\n- **Pass 2 (tuning alone) recovered most of the gap** — the power envelope change from ~60W to ~130W and GPU clocks from ~2000 MHz to ~2850 MHz was the dominant factor.\n- **Pass 3 (IOMMU=off) added the remaining ~7-15%** — consistent with the project’s stated 13-16% IOMMU overhead.\n\n## \n\n### \n\n**Symptoms**: Container starts but immediately exits. Logs show `HIP /src/halogen/src/flash_ops.h:7060: no ROCm-capable device is detected`.\n\n**Causes**:\n\n1. **Rootless Podman** — Device nodes are remapped to`nobody:nogroup` . Use`sudo podman` with`--group-add <render_GID>` .\n2. **User not in render group** — Add your user:`sudo usermod -aG render,video $USER` then log out and back in.\n3. **KFD not initialized** — Check`sudo dmesg | grep kfd` . If empty, the kernel may need`amd_iommu=off` removed or KFD may need`modprobe amdkfd` (though on the AMD AI halo kernel it’s built-in).\n\n### \n\n**Symptoms**: Startup log shows memory budget close to limit, then engine refuses to start.\n\n**Fix**: Reduce KV pool size:\n\n```\n-e HALOGEN_KV_POOL_POSITIONS=524288\n-e HALOGEN_CTX=524288\n```\n\nThis halves the KV cache from 1M to 524K positions, freeing ~14 GiB.\n\n### \n\n**Symptoms**: OOM during weight loading or KV pool reservation.\n\n**Fix**: Reduce `HALOGEN_MAX_TOK` (working memory):\n\n```\n-e HALOGEN_MAX_TOK=16384\n```\n\nThe engine auto-caps this at 16384 for 1M context anyway, but setting it explicitly avoids the warning.\n\n### \n\n**If** a future kernel update changes the KFD behavior with `amd_iommu=off`, rollback:\n\n```\nsudo cp /efi/loader/entries/...conf.bak-iommu /efi/loader/entries/...conf\nsudo reboot\n```\n\nOn kernel 6.18.44+rex+5-amd64 with ROCm 7.14, KFD works correctly via GART fallback. The vendor’s reference machine runs the same configuration.\n\n### \n\nsystemd-boot sorts entries by `sort-key`, then by version. If you edited the 6.18.35 entry but the system booted 6.18.44, apply the edit to the **actually-booted** entry. Check with `bootctl status | grep \"Current Entry\"`. I’m just noting this here because I swear I remember this behaving differently in the past, but I wouldn’t expect the Debian maintainers to have changed this. Maybe Limine on CachyOS is overwriting my memories lol\n\n## \n\n### \n\n```\n# Restore the pre-iommu backup\nsudo cp /efi/loader/entries/amd-ryzen-ai-developer-platform-6.18.44+rex+5-amd64.conf.bak-iommu \\\n       /efi/loader/entries/amd-ryzen-ai-developer-platform-6.18.44+rex+5-amd64.conf\nsudo reboot\n```\n\n### \n\n```\n   sudo -S sh -c 'echo balanced > /sys/firmware/acpi/platform_profile'\nfor f in /sys/devices/system/cpu/cpufreq/policy*/scaling_governor; do\n    sudo -S sh -c \"echo powersave > \\\"$f\\\"\"\ndone\n```\n\n### \n\n```\nsudo podman stop halogen-flash-server\nsudo podman rm halogen-flash-server\n# Then re-run the stock launch command from Section 2\n```\n\n### \n\n``` python\n# Health check\ncurl -s http://localhost:8731/health | python3 -c \"import sys,json; d=json.load(sys.stdin); print('OK' if d['status']=='ok' else 'FAIL')\"\n\n# Quick chat test\ncurl -s -X POST http://localhost:8731/v1/chat/completions \\\n  -H 'Content-Type: application/json' \\\n  -d '{\"model\":\"halogen-qwen3.8-flash-next\",\"messages\":[{\"role\":\"user\",\"content\":\"Say hello in one word.\"}],\"max_completion_tokens\":10,\"temperature\":0,\"enable_thinking\":false}'\n\n# Expected: {\"choices\":[{\"message\":{\"content\":\"Hello\"}}]}\n```\n\n## \n\n- **Peonist AI** for the[halogen-flash-server](https://github.com/peonist-ai/halogen-flash-server) project and the[Qwen3.8-Flash-Next](https://huggingface.co/peonist-ai/halogen-qwen3.8-flash-next) model weights\n- **AMD** for the Strix Halo platform and ROCm 7.14\n- **Qwen team (Alibaba)** for the underlying Qwen3.8 architecture\n\nI am genuinely very impressed with *how fast* the software ecosystem is improving. Halogen apparently cutting away all the cruft has led to more coherence and performance than I would have thought possible on the Strix Halo platform. 45t/s at 1M!", "url": "https://wpnews.pro/news/ryzen-ai-halo-halogen-server-testing-notes", "canonical_source": "https://forum.level1techs.com/t/ryzen-ai-halo-halogen-server-testing-notes/257455#post_2", "published_at": "2026-09-30 18:14:11+00:00", "updated_at": "2026-09-30 18:18:54.406419+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-tools", "mlops"], "entities": ["Halogen Flash Server", "Qwen3.8-Flash-Next", "AMD Strix Halo", "AMD Ryzen AI MAX+ 395", "Radeon 8060S", "ASUS ROG Flow Z13", "Debian 13", "ROCm 7.14"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ryzen-ai-halo-halogen-server-testing-notes", "markdown": "https://wpnews.pro/news/ryzen-ai-halo-halogen-server-testing-notes.md", "text": "https://wpnews.pro/news/ryzen-ai-halo-halogen-server-testing-notes.txt", "jsonld": "https://wpnews.pro/news/ryzen-ai-halo-halogen-server-testing-notes.jsonld"}}