Testing Radiance on 2-4-8 R9700s (and R9600Ds): Some Qualitative Breakout Fun Five of six Qwen3.8-family model runs on AMD RDNA4 GPUs with the Radiance inference engine produced a playable Breakout game from a single 6,567-token prompt with no harness or retry loop, according to a Level1Techs writeup by author StillDeadcode. The tests ran on a 2x ASUS R9700 system and an 8x ASROCK R9600D server, with a 262,144-token budget; the MXFP4 dense model's output needed a two-line fix because its JavaScript never added the `.visible` class that CSS defined for showing game screens. Throughput ranged from 250 to 324 aggregate tokens per second at concurrency 8, with the MXFP4 dense model reaching 678 tok/s peak aggregate and Flash-Next MoE hitting 1,103 tok/s at concurrency 16. This writeup covers the qualitative side of testing Qwen3.8-family models on AMD RDNA4 GPUs with the Radiance inference engine. Our earlier posts covered the performance numbers and setup instructions. This one is about what you can actually expect when you use these models – how they behave, where they break, what the output looks like, and what you need to know about reasoning settings. If you want the how-to guides and benchmark methodology, check out our other writeups: For testing, we used the 2x ASUS R9700 system and our 8x ASROCK R9600D server system. btw, errybody welcome @StillDeadcode https://forum.level1techs.com/u/stilldeadcode , the author We used a 6,567-token programming prompt that asks the model to build a complete Breakout game in HTML/CSS/JavaScript, with a built-in 7-test verification framework. The prompt specifies either a single-file or three-file layout and includes detailed requirements for game mechanics, state management, and event-driven architecture. This is a one-shot test. No harness so the verification framework doesn’t run in this test – more on that in a bit , no retry loop, no feedback cycle. The model gets the prompt, generates until it hits the token budget or emits an EOS token, and whatever it produces is what we evaluate. This is deliberately harder than a harness-based evaluation – a harness would catch the “start button does nothing” class of bug and prompt the model to fix it. We wanted to see what falls out raw before passing to a harness that can run the verification criteria. The token budget was 262,144 the model’s full native context . Every run had enough room to finish. Before the qualitative results, here are the raw throughput numbers for context. | Model | Aggregate tok/s conc=8 | Single-stream sustained | Breakout tok/s | |---|---|---|---| | FP8 dense Swift1.5 27B | 324 | 179 tok/s | 147 | | MXFP4 dense Qwen3.8-27B | 322 | – | 242 | | Swift1.5 Flash-Next MoE | 274 | 175 tok/s | 178 low effort | | Regular Flash-Next MoE | 250 | 162 tok/s | 171 | | Model | Single-stream sustained | Peak aggregate | |---|---|---| | Swift1.5 FP8 dense | 124 tok/s | 586 @ conc8 | | Qwen3.8-27B FP8 dense | 130 tok/s | 596 @ conc8 | | Qwen3.8-27B MXFP4 dense | – | 678 @ conc8 | | Flash-Next MoE | 137 tok/s | 1103 @ conc16 | The dense models are faster in single-stream throughput because they have less total weight to stream per token. Flash-Next compensates at higher concurrency because its MoE architecture activates only 6B parameters per token, leaving more VRAM for KV cache and batch slots. Five out of six models produced a playable game on the first shot. No harness, no retry, no code review loop. The model read a 6,567-token specification and produced a complete, runnable Breakout clone. | Model | Finish reason | Tokens generated | Output size | Game playable? | |---|---|---|---|---| | FP8 dense Swift1.5 | stop | 58,416 | 61 KB | Yes | | MXFP4 dense | stop | 67,988 | 65 KB | Yes after two-line fix | | Swift1.5 Flash-Next low effort | stop | 20,735 | 74 KB | Yes | | Swift1.5 Flash-Next medium effort | stop | 58,178 | 80 KB | Yes | | Swift1.5 Flash-Next xhigh, pp=2 | stop | 86,377 | 88 KB | Yes | | Regular Flash-Next | stop | 78,757 | 64 KB | Yes | All six runs produced the full three-file deliverable index.html, style.css, script.js with proper HTML structure, canvas rendering, game state management, and a verification framework. None of them truncated mid-code or burned their budget on planning prose with one exception, discussed below . The MXFP4 dense model’s output had a subtle but fatal bug: the CSS defined .screen.visible { display: flex; } to show game screens, and the JavaScript dispatched a phase-change custom event when the game state changed – but nothing ever added the .visible class . The event handler updated text content scores, level numbers but never toggled screen visibility. The result: the game loaded without errors, the canvas rendered, the state machine worked internally – but the start button was invisible because every screen had display: none . Clicking where the button should be did nothing. The fix was two lines of JavaScript. We added screen visibility toggling to the phase-change event handler hide all screens, show the one matching the new phase and an initial call in init to show the menu screen on load. The full diff is included in the artifact tarball MANUAL FIX.diff . Here’s the thing: if this had been run through a harness – even a simple one that loads the page, checks for a visible start button, and reports “button not found” – the model would have had the feedback it needed to fix the bug in one retry. The bug is exactly the kind a code review or automated test catches immediately. In a one-shot test with no feedback loop, it’s fatal. The most interesting failure was Swift1.5 Flash-Next at its default reasoning effort xhigh . The model generated 255,555 tokens – nearly a million characters of reasoning content – without producing a single line of actual code. The reasoning content was self-referential and repetitive: the model kept listing things it “needed to ensure” about the game state, in patterns like: “Need ensure no use of state.ghostClearOnGhostClearOnLevel? no. Need ensure no use of state.ghostClearOnGhostClearOnLives? no. Need ensure no use of state.ghostClearOnGhostClearOnMiss? no.” This continued for hundreds of thousands of tokens. The model was stuck in a loop where each “need to check X” generated more “need to check Y” thoughts, never converging on actual output, but also several dozens of lines before a repeat would occur. The fix was presence penalty: 2.0 – a setting the Qwen model card explicitly documents as reducing “endless repetition.” With presence penalty at 2.0 and reasoning effort at xhigh, the model produced 86,377 tokens of reasoning still a lot, but bounded followed by a complete, working game. Qwen is a “nervous test taker” ? the model wasn’t broken. It was doing exactly what we told it to do – think as deeply as possible xhigh – and it thought itself into a corner. The repetition penalty we needed was literally documented on the model card. We just hadn’t read that far before running the test. When we set reasoning effort to “low,” the model produced a clean game in 20,735 tokens with only 7,380 tokens of reasoning. It stopped overthinking because we stopped telling it to overthink. | Setting | tok/s | Reasoning tokens | Content tokens | Result | |---|---|---|---|---| | low | 178 | 7,380 | 20,735 | Clean code, 117 seconds | | medium | 137 | 135,581 | 58,178 | Clean code, 424 seconds | | xhigh + pp=2 | 110 | 247,512 | 86,377 | Clean code, 786 seconds | | xhigh + pp=0 | 184 | 921,676 | 0 | Reasoning loop, no code | The Qwen model card documents reasoning effort values of xhigh default , medium, and low. It also notes that “in multi-turn agentic tasks, lower reasoning effort doesn’t always reduce total completion time” – which is true for chat, but for code generation the opposite is clearly the case. Lower effort meant faster completion and better output, because the model spent its token budget on code instead of self-referential reasoning. There is a great qualitative aspect to this, so here’s the output from each of the runs for science. Each artifact tarball includes a screenshot.png from the browser showing what the game actually looks like when loaded… The FP8 dense model’s output had an odd bug I didn’t catch at first: buildLevel referenced CONFIG.grid.rowColors row , but rowColors was defined at the CONFIG level, not inside CONFIG.grid . This caused a TypeError: Cannot read properties of undefined reading '0' but I wasn’t paying attention to the console. The fix was changing one word : g.rowColors row to CONFIG.rowColors row . Again, this is exactly the kind of bug a harness would catch immediately – the page throws an error on load, the harness reports it, the model fixes it. Each breakout clone has minor visual glitches, edge-case bugs, and polish issues that we didn’t catalog. The games are playable but not production-quality. Some examples: Our general impression: the less quantized the model, the better the result. The FP8 dense model 8-bit trunk produced the cleanest code with the fewest bugs. The MXFP4 model 4-bit experts had the screen visibility bug. The Flash-Next models 4-bit experts + int8 trunk produced working games but with more reasoning overhead. This is a sample size of one prompt, but I thought it was a good cross-section of my experiments on AMD hardware with Radiance. And a good harness makes the result that much better. Every bug we found – the MXFP4 screen visibility, the fp8 dense rowColors reference, the Flash-Next reasoning loop – would have been caught and fixed in one or two retry cycles with even a minimal evaluation harness. Every breakout run’s complete output is packaged for evaluation. Each tarball contains the extracted game files index.html, style.css, script.js , a breakout.html you can double-click to play, the raw model output for reference, and metadata about the model, settings, and results. fp8 dense.tar.gz https://forum.level1techs.com/uploads/short-url/9wGykZ70XfNR4EfJk4k9hbUjLZy.gz 63.3 KB mxfp4 dense.tar.gz https://forum.level1techs.com/uploads/short-url/z8ozQNsP471SyWebWRaewh0CT7w.gz 69.0 KB regular fn.tar.gz https://forum.level1techs.com/uploads/short-url/sf156Q3fHlV9ns1pP9y3bt5J4Uo.gz 58.6 KB swift fn low.tar.gz https://forum.level1techs.com/uploads/short-url/qN6x9KzVwU8TDZL67zCuLRvhKR2.gz 62.7 KB swift fn medium.tar.gz https://forum.level1techs.com/uploads/short-url/nJJBb9cWkfKNYEKo39zQw7RvyIa.gz 113.6 KB swift fn xhigh.tar.gz https://forum.level1techs.com/uploads/short-url/cMGazTbOQqe3qiYWYMqSHagFnEy.gz 324.8 KB | Artifact | Model | Size | Notes | |---|---|---|---| | fp8 dense.tar.gz | Swift1.5-Qwen3.8-27B FP8 | 65 KB | Includes the rowColors fix | | mxfp4 dense.tar.gz | Qwen3.8-27B MXFP4 | 71 KB | Includes the screen visibility fix + diff | | swift fn low.tar.gz | Swift1.5 Flash-Next low effort | 64 KB | Cleanest run, no fixes needed | | swift fn medium.tar.gz | Swift1.5 Flash-Next medium effort | 116 KB | Includes reasoning trace | | swift fn xhigh.tar.gz | Swift1.5 Flash-Next xhigh, pp=2 | 333 KB | Includes 248 KB reasoning trace | | regular fn.tar.gz | Qwen3.8-Flash-Next FP8+IQ4R | 60 KB | No fixes needed | breakout prompt.txt https://forum.level1techs.com/uploads/short-url/s2pXeLy49GDuU4RZP2GbBjO2ei0.txt 28.9 KB Download the artifacts, extract, and double-click breakout.html to evaluate the games yourself. The MXFP4 tarball includes both breakout.html the original, broken output and breakout with patch.html the fixed version so you can see the difference. Radiance https://codeberg.org/StillDeadcode/radiance is a standalone C++/HIP inference engine for AMD RDNA4 GPUs, written from scratch by @StillDeadcode https://forum.level1techs.com/u/stilldeadcode . It started as vllm-radiance a patched vLLM fork with hand-written RDNA4 kernels and evolved into a fully independent engine with a plugin architecture. The Docker image is 423 MB. It supports FP8, MXFP4, and bf16 quantization, speculative decoding MTP and DFlash2 , tensor parallelism up to TP4, and expert tiering for MoE models. For setup instructions, benchmark methodology, and performance tuning, see our other writeups elsewhere on this forum. The short version: pull stilldeadcode/radiance:latest , download a .rad model container from HuggingFace, and run with --tp 4 --tp-wire wht6 --p2p off on an 8-GPU system. Six models, one prompt, no harness. The throughput numbers are strong: 147-242 tok/s for single-stream code generation on consumer AMD Radeon RDNA GPUs, with aggregate throughput up to 1103 tok/s at concurrency 2x R9700s don’t wanna spoil the peak numbers from 8x R9600s just yet haha . The qualitative results are encouraging: these models can build complete, runnable applications from a detailed specification in a single pass.