Ryzen AI Halo: Halogen Server Testing Notes A community tester published deployment notes for the Halogen Flash Server, an optimized server for running Qwen3.8-Flash-Next on AMD Strix Halo (gfx1151) hardware, reporting a roughly 10% performance uplift from the amd_iommu=off kernel parameter. The notes cover a real-world deployment on an AMD Ryzen AI MAX+ 395 mini-PC with 128 GB unified memory, using Debian 13, Podman, and ROCm 7.14, and walk through stock launch, power tuning, YaRN 1M context extension, and container management. The model weights and n-gram table total about 111 GB, split between a 63 GiB main checkpoint, a 48 GiB n-gram lookup table for speculative decoding, and an optional 857 MiB vision sidecar. I dunno that this will become a video, theres so many videos but I wanted to share this with the community as its my notes from experimenting with the “Halogen” engine on the Strix Halo platform. Maybe I’ll tweet about it. What’s Halogen? It’s this sort of interesting optimized server for running qwen 3.8 flash next on Strix halo. You should read about the project, it’s interesting. This was tested on several Strix Halo platforms including the ASUS ROG Flow Z13. Original project : github.com/peonist-ai/halogen-flash-server https://github.com/peonist-ai/halogen-flash-server — the fastest way to run Qwen3.8-Flash-Next on AMD Strix Halo gfx1151 . Model : peonist-ai/halogen-qwen3.8-flash-next https://huggingface.co/peonist-ai/halogen-qwen3.8-flash-next on Hugging Face. Platform : AMD Ryzen AI MAX+ 395 Radeon 8060S, 128 GB unified memory . This guide covers a real-world deployment on a Strix Halo mini-PC. Normally I use Ubuntu LTS with Docker, but I decided to try Podman on Debian 13 this time around. We walk through the stock launch, power tuning, YaRN 1M context extension, and the amd iommu=off kernel parameter that’s good for about +10% perf uplift. Each step includes the commands, the reasoning, and the performance impact. 1. Prerequisites & Hardware 1-prerequisites--hardware 2. Stock Launch Reference 2-stock-launch-reference 3. Power Tuning: Platform Profile & CPU Governor 3-power-tuning--platform-profile--cpu-governor 4. YaRN 1M Context Extension 4-yarn-1m-context-extension 5. amd iommu=off – The Secret Sauce 5-amd iommuoff--the-secret-sauce 6. Container Management 6-container-management 7. Performance Results 7-performance-results 8. Troubleshooting 8-troubleshooting 9. Appendix: Rollback & Recovery 9-appendix-rollback--recovery AMD Ryzen AI Halo GMKtec ROG Flow Z13 Framework 13 These systems are all bananas. The software has made the story. This should all work on AMD’s Gorgon Halo as well – 400 series with 192gb ram. These all had 128gb LPDDR5 unified memory. | Component | Detail | | CPU/APU | AMD Ryzen AI MAX+ 395 16C/32T | | GPU | Radeon 8060S gfx1151, integrated, 128 GB unified memory | | RAM | 128 GB LPDDR5X shared CPU/GPU pool | | Storage | NVMe SSD | | Platform | Mini-PC ASUS ROG Flow Z13 | - OS : Debian 13 “Trixie” or any recent Linux with ROCm 7.14 support - Container runtime : Podman Docker works too; adjust --group-add accordingly - ROCm : 7.14 shipped with the halogen-flash-server image - Kernel : 6.18.44+rex+5-amd64 vendor kernel from AMD on the AI Halo, but CachyOS kernel tested and works too. Download the model weights and n-gram table from Hugging Face ~111 GB total : Install huggingface-cli if needed pip install huggingface-hub Download model huggingface-cli download peonist-ai/halogen-qwen3.8-flash-next \ --local-dir ~/halogen-models ^ strictly speaking this isn’t needed as the halogen github link, at first run, downloads the model. I like to do this to ensure that “extra” downloads across test runs do not happen. Save my precious bandwidth. Expected files: | File | Size | Purpose | | qwen38-flash-next-v2.hgn | 63 GiB | Main checkpoint weights | | qwen38-flash-next-ngram.hgn | 48 GiB | N-gram lookup table for speculative decoding | | qwen38-flash-next-vision.hgn | 857 MiB | Vision sidecar optional, not loaded by default | | tokenizer/ | — | Tokenizer files | If you aren’t read in on the whole N-gram thing, this is a “new” thing where part of the model is designed to run from dram at dram speeds. This is a unified memory platform so it manifests as a speed benefit, but the N-gram model work should be exciting to everyone since it promises decent performance on “normal” systems that have fast vram faster than LPDDR5 and slower system memory. This is the baseline from the GitHub README https://github.com/peonist-ai/halogen-flash-server . Run the container as-is, no tuning: podman run -d --name halogen-flash-server \ -p 8731:8731 \ --device /dev/kfd --device /dev/dri \ --group-add keep-groups \ --ipc=host \ --ulimit memlock=-1:-1 \ -v ~/halogen-models:/models \ ghcr.io/peonist-ai/halogen-flash-server:0.15.0 The server listens on port 8731. Check health: curl -s http://localhost:8731/health | python3 -m json.tool Key observations at stock: - Platform profile: balanced - CPU governor: powersave - GPU clocks: ~2000 MHz under load - GPU power: ~60W - Context: 262,144 tokens native, no YaRN - Output cap: 65,536 tokens The stock balanced profile and powersave governor leave a lot of performance on the table. I suspect some of the benchmarks posted on github on strix halo were run with a conservative CPU governor. The GPU can hit 2900 MHz and 130W, but the power envelope needs to be opened up. The GMKtek has been tuned for 150W, too, but I have omitted the results from this writeup to keep it from being confusing it was only worth about 2% more, at best, performance . cat /sys/firmware/acpi/platform profile → balanced cat /sys/devices/system/cpu/cpufreq/policy /scaling governor | sort -u → powersave cat /sys/class/drm/card0/device/pp dpm sclk 0: 600Mhz 1: 1100Mhz 2: 2900Mhz Set platform profile to performance sudo -S sh -c 'echo performance /sys/firmware/acpi/platform profile' Set all CPU governors to performance for f in /sys/devices/system/cpu/cpufreq/policy /scaling governor; do sudo -S sh -c "echo performance \"$f\"" done cat /sys/firmware/acpi/platform profile → performance cat /sys/devices/system/cpu/cpufreq/policy0/scaling governor → performance | Parameter | Before | After | | Platform profile | balanced | performance | | CPU governor | powersave | performance | | GPU clock under load | ~2000 MHz | ~2850 MHz | | GPU power under load | ~60W | ~130W | sudo -S sh -c 'echo balanced /sys/firmware/acpi/platform profile' for f in /sys/devices/system/cpu/cpufreq/policy /scaling governor; do sudo -S sh -c "echo powersave \"$f\"" done Note : These settings are ephemeral — they reset on reboot. See the Post-Reboot Automation post-reboot-automation section to persist them. The model’s native context is 262,144 tokens. With YaRN Yet another RoPE extensioN at factor 4, we extend this to 1,048,576 tokens ~1M . This requires a larger KV cache pool — about 28.8 GiB of the 128 GiB unified memory. podman run -d --name halogen-flash-server \ -p 8731:8731 \ --device /dev/kfd --device /dev/dri \ --group-add keep-groups \ --ipc=host \ --ulimit memlock=-1:-1 \ -e HALOGEN ROPE YARN=4 \ -e HALOGEN CTX=1048576 \ -e HALOGEN MAX THINKING TOKENS=262144 \ -e HALOGEN MAX TOKENS DEFAULT=393216 \ -e HALOGEN MAX TOKENS CAP=393216 \ -v ~/halogen-models:/models \ ghcr.io/peonist-ai/halogen-flash-server:0.15.0 | Variable | Value | Purpose | | HALOGEN ROPE YARN=4 | 4 | YaRN scaling factor. 4 × 262K native = 1,048,576 context | | HALOGEN CTX=1048576 | 1,048,576 | Max context length KV pool positions | | HALOGEN MAX THINKING TOKENS=262144 | 262,144 | Reasoning budget cap 256K tokens for thinking | | HALOGEN MAX TOKENS DEFAULT=393216 | 393,216 | Default total output budget thinking + answer | | HALOGEN MAX TOKENS CAP=393216 | 393,216 | Hard cap; requests above get 400 error | python curl -s http://localhost:8731/health | python3 -c " import sys, json d = json.load sys.stdin print 'Context:', d 'context' print 'YaRN:', d.get 'rope scaling' print 'Max tokens default:', d.get 'max tokens default' print 'Max thinking tokens:', d.get 'max thinking tokens default' " Expected output: Context: 1048576 YaRN: {'type': 'yarn', 'factor': 4.0, 'original context': 262144} Max tokens default: 393216 Max thinking tokens: 262144 rope: static YaRN, factor 4 over the native 262144 attention scale 1.138629 startup 3.6 s KV pool reserved: 1048576 positions about 28.8 GiB startup 3.7 s memory: 62.1 GiB of weights locked in RAM, 28.8 GiB of KV pool, 8.1 GiB of working memory, 99.0 GiB in all startup 3.7 s host memory left for everything else: ~20 GiB With YaRN 1M enabled, the server uses ~99 GiB of the 128 GiB pool: - 62 GiB — model weights pinned in RAM - 29 GiB — KV cache 1,048,576 positions - 8 GiB — working memory - ~20 GiB — remaining for the OS and other services If you’re tight on memory, you can halve the KV pool: -e HALOGEN KV POOL POSITIONS=524288 -e HALOGEN CTX=524288 This reduces context to 524K but frees ~14 GiB. The single biggest performance unlock on Strix Halo is disabling the IOMMU. On this platform, the IOMMU adds overhead to GPU memory access that visibly impacts prefill throughput. The project’s reference machine https://github.com/peonist-ai/halogen-flash-server runs with amd iommu=off as part of its kernel command line, and the README notes it’s worth 13–16% of prefill performance . Note : The reference machine also uses additional GPU-tuning kernel parameters: amdgpu.vm update mode=0 amdgpu.noretry=0 amdgpu.gttsize=126976 ttm.pages limit=32505856 amdgpu.sg display=0 . We did NOT apply these in our testing — our results come from amd iommu=off alone, combined with the userspace power tuning from Section 3. The system uses systemd-boot not GRUB . Boot entries live on the EFI System Partition at /efi/loader/entries/ . Step 1: Identify the current boot entry bootctl status | grep "Current Entry" → amd-ryzen-ai-developer-platform-6.18.44+rex+5-amd64.conf Step 2: Back up the entry sudo cp /efi/loader/entries/amd-ryzen-ai-developer-platform-6.18.44+rex+5-amd64.conf \ /efi/loader/entries/amd-ryzen-ai-developer-platform-6.18.44+rex+5-amd64.conf.bak-iommu Step 3: Add amd iommu=off to the kernel command line sudo sed -i 's/^options /options amd iommu=off /' \ /efi/loader/entries/amd-ryzen-ai-developer-platform-6.18.44+rex+5-amd64.conf Step 4: Verify the edit cat /efi/loader/entries/amd-ryzen-ai-developer-platform-6.18.44+rex+5-amd64.conf Should show: options amd iommu=off root=UUID=... splash quiet loglevel=3 Step 5: Reboot sudo systemctl reboot Step 6: Verify after reboot cat /proc/cmdline | tr ' ' '\n' | grep iommu → amd iommu=off sudo dmesg | grep -c AMD-Vi → 0 IOMMU is disabled sudo dmesg | grep -i kfd → kfd kfd: amdgpu: added device 1002:1586 KFD still works Yes. On this kernel 6.18.44+rex+5-amd64 , the AMD KFD Kernel Fusion Driver falls back to GART-based initialization when the IOMMU is disabled. You’ll see this in dmesg: kfd kfd: amdgpu: Allocated 3969056 bytes on gart kfd kfd: amdgpu: Total number of KFD nodes to be created: 1 kfd kfd: amdgpu: added device 1002:1586 ROCm 7.14 works normally. The container needs --group-add with the render group’s GID typically 992 because rootless Podman remaps device ownership, but that’s a container-runtime detail, not an IOMMU issue. If the system fails to boot after adding amd iommu=off , use the boot menu to select the backup entry or use a recovery USB to restore the backup: From a recovery shell: sudo cp /efi/loader/entries/...conf.bak-iommu /efi/loader/entries/...conf On this Debian 13 system, rootless Podman remaps device node ownership inside the container. The /dev/kfd and /dev/dri/renderD128 devices show up as owned by nobody:nogroup , which breaks ROCm’s permission check. Fix : Run the container under sudo podman with the render group’s GID: Find the render GID grep ^render /etc/group | cut -d: -f3 → 992 Launch with sudo sudo podman run -d --name halogen-flash-server \ -p 8731:8731 \ --device /dev/kfd --device /dev/dri \ --group-add 992 \ --ipc=host \ --ulimit memlock=-1:-1 \ -e HALOGEN ROPE YARN=4 \ -e HALOGEN CTX=1048576 \ -e HALOGEN MAX THINKING TOKENS=262144 \ -e HALOGEN MAX TOKENS DEFAULT=393216 \ -e HALOGEN MAX TOKENS CAP=393216 \ -v /home/w/halogen-models:/models \ ghcr.io/peonist-ai/halogen-flash-server:0.15.0 Save this as ~/halogen-podman and chmod +x : bash /bin/bash PASS=your sudo password here ACTION="$1" shift case "$ACTION" in start|stop|restart|logs|rm echo "$PASS" | sudo -S podman "$ACTION" halogen-flash-server "$@" ;; ps echo "$PASS" | sudo -S podman ps -a --filter name=halogen-flash-server "$@" ;; health echo "$PASS" | sudo -S curl -s http://localhost:8731/health "$@" ;; run echo "$PASS" | sudo -S podman run -d --name halogen-flash-server \ -p 8731:8731 \ --device /dev/kfd --device /dev/dri \ --group-add 992 \ --ipc=host \ --ulimit memlock=-1:-1 \ -e HALOGEN ROPE YARN=4 \ -e HALOGEN CTX=1048576 \ -e HALOGEN MAX THINKING TOKENS=262144 \ -e HALOGEN MAX TOKENS DEFAULT=393216 \ -e HALOGEN MAX TOKENS CAP=393216 \ -v /home/w/halogen-models:/models \ ghcr.io/peonist-ai/halogen-flash-server:0.15.0 ;; echo "Usage: halogen-podman {start|stop|restart|logs|rm|ps|health|run}" ;; esac Usage: ./halogen-podman start Start the container ./halogen-podman stop Stop it ./halogen-podman restart Restart it ./halogen-podman logs View logs ./halogen-podman health Check health endpoint ./halogen-podman ps Container status ./halogen-podman run Create a fresh container There is probably a better way to do this. This kind of thing isn’t necessary with Docker, or I might need more XP with podman to understand what I’ve missed. I suspect that because when you’re in the docker group, you’ve basically got root, and this is a different architectural choice in podman to prevent that. But I need the hardware. So this is just making some notes for myself about that. After every reboot, you need to: 1. Re-apply the performance profile + governor 2. Start the halogen container A simple systemd oneshot service can handle this kind of thing, though. Create /etc/systemd/system/halogen-tune.service : Unit Description=Halogen power tuning After=multi-user.target Service Type=oneshot ExecStart=/usr/local/bin/halogen-tune.sh RemainAfterExit=yes Install WantedBy=multi-user.target And /usr/local/bin/halogen-tune.sh : bash /bin/bash echo performance /sys/firmware/acpi/platform profile for f in /sys/devices/system/cpu/cpufreq/policy /scaling governor; do echo performance "$f" done sudo systemctl enable halogen-tune.service Then add the container start to your crontab @reboot /home/w/halogen-podman start or another systemd unit. Different distros handle this in different ways. I’m getting a bit rusty on Debian and there may be a less brute-force way to do this. But this is good practice for you sysadmin-in-training folks out there. We used the built-in sweep command, which measures end-to-end HTTP request latency including prefill and decode: Prefill benchmarks 3 repetitions sudo podman run --rm ... halogen-flash-server:0.15.0 sweep -p 8192,32768 -n 128 -r 3 sudo podman run --rm ... halogen-flash-server:0.15.0 sweep -p 131072 -n 128 -r 1 Decode benchmark 10 prompt shapes, 1 repetition sudo podman run --rm ... halogen-flash-server:0.15.0 bench mtp 256 low 1 All benchmarks use the MTP Multi-Token Prediction drafter, which is the default and fastest mode. | Test | Ref 0.14.1 | Baseline 0.14.0 | Pass 1 stock | Pass 2 tuned | Pass 3 tuned+IOMMU | | pp8192 | 1,584 t/s | 1,268 t/s | 1,350 t/s | 1,679 t/s | 1,768 t/s | | pp32768 | 1,567 t/s | 1,442 t/s | 1,216 t/s | 1,469 t/s | 1,695 t/s | | pp131072 | 1,517 t/s | 1,390 t/s | 1,171 t/s | 1,426 t/s | 1,621 t/s | | tg128 MTP | 46.0 t/s | 39.9 t/s | 48.1 t/s | 52.0 t/s | 52.2 t/s | | mtp 256 low 1 | — | — | — | — | 53.4 t/s | pp = prefill prompt processing , tg = token generation. All values in tokens/second. Higher is better. Reference and Baseline columns from the project README https://github.com/peonist-ai/halogen-flash-server . | Pass | Changes | | Ref 0.14.1 | Project’s reference machine – ROCm 7.14, IOMMU=off, additional GPU kernel params, ~85W sustained | | Baseline 0.14.0 | Same machine, prior software version for comparison | | Pass 1 | Our stock baseline – balanced profile, powersave governor, IOMMU on | | Pass 2 | Platform profile → performance , CPU governor → performance on all 32 cores | | Pass 3 | Pass 2 + amd iommu=off kernel parameter + reboot | Most real-world usage will be at 256K context or less – the model’s native length. The YaRN 1M config is available for deep-context tasks, but 256K covers the vast majority of chat, coding, and analysis workloads. Here’s how performance compares: | Metric | 256K Context native | 1M Context YaRN 4x | | KV pool size | ~7 GiB | ~29 GiB | | Total memory used | ~77 GiB | ~99 GiB | | Host memory free | ~42 GiB | ~20 GiB | | Prefill @ 8K | ~1,770 t/s | ~1,770 t/s | | Prefill @ 32K | ~1,700 t/s | ~1,700 t/s | | Prefill @ 131K | ~1,620 t/s | ~1,620 t/s | | Decode MTP, short ctx | ~53 t/s | ~53 t/s | | Decode MTP, deep ctx ~260K | — | ~45 t/s | The prefill and short-context decode numbers are nearly identical between configs – the YaRN scaling doesn’t add meaningful overhead for prompt processing or short generations. The difference appears at deep context: the README reports 45.0 tok/s decode at 258K context with the 1M config, vs 46.0 tok/s at 32K context. The 1M config also enables prompt caching across very long sessions the follow-up turn at 100K context is ~2s . Recommendation : Run with the 1M config by default. The memory cost ~22 GiB extra for the larger KV pool is worth the flexibility, and performance at 256K seems basically identical. Only drop to 524K or 256K if you’re running other memory-hungry services alongside the server. | Metric | Stock Pass 1 | Tuned Pass 2 | Tuned+IOMMU Pass 3 | Ref Machine | | GPU clock under load | ~2000 MHz | ~2850 MHz | ~2900 MHz | — | | GPU power under load | ~60W | ~130W | ~140W | ~85W | | Idle clock | 600 MHz | 600 MHz | 600 MHz | — | | Idle power | ~5W | ~5W | ~5W | — | The GPU hits its 2900 MHz ceiling and 140W power limit 160w on the GMKtek after tuning. The reference machine runs at ~85W sustained – our higher power draw is expected given the performance profile the reference likely uses a tuned balanced profile with IOMMU=off and the additional GPU kernel parameters . The remaining small gap vs reference at 131K prefill is likely memory-bandwidth bound rather than clock-limited. pp8192: ████████████████████░░░░░░░░░░ 1,768 t/s +12% vs ref 1,584 pp32768: █████████████████████░░░░░░░░░░ 1,695 t/s +8% vs ref 1,567 pp131072: ████████████████████████░░░░░░░░ 1,621 t/s +7% vs ref 1,517 Decode: ██████████████████████████░░░░░░ 53.4 t/s +16% vs ref 46.0 - Pass 3 tuned + IOMMU=off exceeds the reference at every prefill size, despite not using the additional GPU kernel parameters amdgpu.vm update mode=0 , etc. that the reference machine employs. This suggests the userspace power tuning performance profile + governor is doing significant work beyond just IOMMU=off. - Pass 1 was below the 0.14.0 baseline — our stock configuration IOMMU on, balanced profile was leaving ~20% on the table. - Pass 2 tuning alone recovered most of the gap — the power envelope change from ~60W to ~130W and GPU clocks from ~2000 MHz to ~2850 MHz was the dominant factor. - Pass 3 IOMMU=off added the remaining ~7-15% — consistent with the project’s stated 13-16% IOMMU overhead. Symptoms : Container starts but immediately exits. Logs show HIP /src/halogen/src/flash ops.h:7060: no ROCm-capable device is detected . Causes : 1. Rootless Podman — Device nodes are remapped to nobody:nogroup . Use sudo podman with --group-add