{"slug": "the-npu-works-the-compiler-doesn-t", "title": "The NPU Works. The Compiler Doesn't.", "summary": "A developer testing an AMD Ryzen AI NPU (~50 TOPS) on a Beelink SER9 found the hardware, driver, and runtime all work — including running passthrough kernels, conv2d, and a ResNet bottleneck block directly on the NPU tiles — but the open-source IREE `amd-aie` compiler backend fails at the final `translate-to-AIE` step for any convolution or matmul-with-bias, lowering only plain matmul. The failure means a YOLOX backbone cannot compile for the NPU because the fundamental convolution op has no code path, a gap the author attributes to operator coverage in the open-source AIE code generator rather than the silicon.", "body_md": "I wanted a cheap, low-power, always-on person detector. Of course, I’m building this on a BrightSign XT5 player (more on that later).  But since a friend said he was using a Beelink SER9 with a AMD Ryzen AI NPU (~50 TOPS) I thought I’d check it out. Can it actually run the detector? Short answer, as of today: **no — and the reason is not the part you’d expect.**\n\nHere’s the twist that makes this worth a blog post. The hardware works. The driver works. The runtime works. I ran tiny kernels *on the NPU tiles* and watched them come back correct. Every layer of the stack passed its test… right up until the very last one. The thing that’s broken isn’t the silicon. It’s the compiler.\n\n## What I was actually trying to do\n\nThe pitch for an NPU is simple: it’s a dedicated inference block that does sustained neural-net math for almost no power. If I want a box watching a few cameras - maybe more than a few - I do *not* want to burn 40–60 watts pinning the Radeon 890M iGPU to do it. The NPU should do that job at a few watts and never break a sweat.\n\nI’ve [written about running YOLOX on an NPU before](https://blog.herlein.com/post/argus-and-the-npu-repo-family/) — that’s exactly what Argus does on BrightSign players, and it’s genuinely boring in the best way. Drop a file on an SD card, plug in a webcam, boot. So I went in expecting the AMD part to be *fiddly*, sure, but fundamentally a solved problem. But hey, why not see if the AMD box can do the job?\n\n## Everything worked… until the last step\n\nLet me be specific, because “it doesn’t work” is a lazy thing to say and I hate reading it in other people’s posts. Here’s what I got working, layer by layer:\n\n- **Kernel + driver.** Ubuntu 26.04 on kernel 7.0 ships the`amdxdna` driver*in-tree* . No DKMS, no out-of-tree module dance. It just shows up at`/dev/accel/accel0` .\n- **Runtime.** Built XRT from source (SHIM-only, since the kernel already has the driver), and`xrt-smi examine` cheerfully reports`RyzenAI-npu4` , arch`aie2p` , 6×8 tiles. The box sees its NPU.\n- **Actual kernels on the actual tiles.** Using IRON/MLIR-AIE I ran a passthrough kernel (~97 µs of NPU time), a`conv2d` , and a ResNet bottleneck block. They ran*on the NPU* and produced correct results verified against PyTorch.**The silicon executes compiled compute. Full stop.**\n- **The model compiler’s front half.** Feeding an ONNX model through IREE + the`amd-aie` backend: it imports the ONNX, lowers it through torch → linalg, and splits it into dispatches. All clean.\n\nAnd then, at the final `translate-to-AIE` step — the part that turns a lowered op into code that actually runs on the tile array — it falls over:\n\n```\nerror: failed to run translation of source executable to target executable\nfor backend #hal.executable.target<\"amd-aie\", \"amdaie-pdi-fb\", ...>\n```\n\nAn arbitrary convolution (16-channel, 3×3, 32×32) fails. A little matmul-with-bias — i.e. a fully-connected layer — fails. The *one* op class the open-source `iree-amd-aie` code generator lowers reliably today is plain **matmul**. Its whole CI is basically built around a `run_matmul_test.sh`.\n\nThink about that for a minute. A convolution is the fundamental building block of every CNN detector on earth. If you can’t lower a single conv onto the tiles, a YOLOX backbone doesn’t have a prayer of compiling. Not because it’s too big. Because op #1 doesn’t have a code path yet.\n\n## This is a compiler problem, not a hardware problem\n\nThis is the whole point, so let me say it plainly: **the gap is in the open-source AIE code generator’s operator coverage, not in the chip.**\n\nLowering a convolution onto a spatial array of AI-engine tiles is genuinely hard — you have to tile the computation, choreograph all the data movement between tiles, and generate per-core kernels. The `amd-aie` backend just doesn’t implement that for the general case yet. It does matmul. Real models interleave conv, matmul, pooling, SiLU, concat, resize, and NMS, and today a single global tiling-pipeline flag has to serve every dispatch. That’s not enough machinery to compile a detector.\n\nAnd yes, there’s an “official” high-level path — ONNX Runtime with the VitisAI execution provider. It’s Windows-only. On Linux the `voe` graph-optimizer wheels are missing and a config-parser bug quietly shoves every op back onto the CPU, so you *think* you’re using the NPU and you’re not. Ultralytics and PyTorch don’t have a `device=npu` either. The Linux high-level story just isn’t there.\n\nDon’t get me wrong — none of this is me dunking on AMD. The hardware is real, the driver landed in-tree, and the runtime foundation is solid. This is an *ecosystem-maturity* gap, and those close over time. It reminds me of something I say about the models all the time: today’s models are the worst you’ll ever use. Same energy here — today’s `iree-amd-aie` is the worst conv coverage it’ll ever have. It only goes one direction.\n\n## So what do you do right now?\n\nSeriously? Go build it on a BrightSign! Stay tuned…\n\nI’ll revisit this the moment `iree-amd-aie` grows a real convolution path — and when I do, you’ll read about it here first.\n\nIf this saved you a weekend, or if you’ve gotten a conv to lower on XDNA2 under Linux and I’m just holding it wrong, [drop me a note on LinkedIn](https://www.linkedin.com/in/gherlein/). I’d genuinely love to be wrong about the timeline.", "url": "https://wpnews.pro/news/the-npu-works-the-compiler-doesn-t", "canonical_source": "https://blog.herlein.com/post/npu-works-compiler-doesnt/", "published_at": "2026-09-29 08:00:01+00:00", "updated_at": "2026-09-29 19:19:28.021591+00:00", "lang": "en", "topics": ["ai-chips", "ai-infrastructure", "computer-vision", "machine-learning", "ai-tools"], "entities": ["AMD", "Ryzen AI NPU", "Beelink SER9", "BrightSign XT5", "IREE", "MLIR-AIE", "XRT", "YOLOX"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/the-npu-works-the-compiler-doesn-t", "markdown": "https://wpnews.pro/news/the-npu-works-the-compiler-doesn-t.md", "text": "https://wpnews.pro/news/the-npu-works-the-compiler-doesn-t.txt", "jsonld": "https://wpnews.pro/news/the-npu-works-the-compiler-doesn-t.jsonld"}}