This writeup covers the qualitative side of testing Qwen3.8-family models on AMD RDNA4 GPUs with the Radiance inference engine. Our earlier posts covered the performance numbers and setup instructions. This one is about what you can actually expect when you use these models – how they behave, where they break, what the output looks like, and what you need to know about reasoning settings.
If you want the how-to guides and benchmark methodology, check out our other writeups:
For testing, we used the 2x ASUS R9700 system and our 8x ASROCK R9600D server system.
( btw, errybody welcome [@StillDeadcode](https://forum.level1techs.com/u/stilldeadcode) , the author )
We used a 6,567-token programming prompt that asks the model to build a complete Breakout game in HTML/CSS/JavaScript, with a built-in 7-test verification framework. The prompt specifies either a single-file or three-file layout and includes detailed requirements for game mechanics, state management, and event-driven architecture.
This is a one-shot test. No harness (so the verification framework doesn’t run in this test – more on that in a bit), no retry loop, no feedback cycle. The model gets the prompt, generates until it hits the token budget or emits an EOS token, and whatever it produces is what we evaluate. This is deliberately harder than a harness-based evaluation – a harness would catch the “start button does nothing” class of bug and prompt the model to fix it. We wanted to see what falls out raw before passing to a harness that can run the verification criteria.
The token budget was 262,144 (the model’s full native context). Every run had enough room to finish.
Before the qualitative results, here are the raw throughput numbers for context.
| Model | Aggregate tok/s (conc=8) | Single-stream sustained | Breakout tok/s |
|---|---|---|---|
| FP8 dense (Swift1.5 27B) | 324 | 179 tok/s | 147 |
| MXFP4 dense (Qwen3.8-27B) | 322 | – | 242 |
| Swift1.5 Flash-Next MoE | 274 | 175 tok/s | 178 (low effort) |
| Regular Flash-Next MoE | 250 | 162 tok/s | 171 |
| Model | Single-stream sustained | Peak aggregate |
|---|---|---|
| Swift1.5 FP8 dense | 124 tok/s | 586 @ conc8 |
| Qwen3.8-27B FP8 dense | 130 tok/s | 596 @ conc8 |
| Qwen3.8-27B MXFP4 dense | – | 678 @ conc8 |
| Flash-Next MoE | 137 tok/s | 1103 @ conc16 |
The dense models are faster in single-stream throughput because they have less total weight to stream per token. Flash-Next compensates at higher concurrency because its MoE architecture activates only 6B parameters per token, leaving more VRAM for KV cache and batch slots.
Five out of six models produced a playable game on the first shot. No harness, no retry, no code review loop. The model read a 6,567-token specification and produced a complete, runnable Breakout clone.
| Model | Finish reason | Tokens generated | Output size | Game playable? |
|---|---|---|---|---|
| FP8 dense (Swift1.5) | stop | 58,416 | 61 KB | Yes |
| MXFP4 dense | stop | 67,988 | 65 KB | Yes (after two-line fix) |
| Swift1.5 Flash-Next (low effort) | stop | 20,735 | 74 KB | Yes |
| Swift1.5 Flash-Next (medium effort) | stop | 58,178 | 80 KB | Yes |
| Swift1.5 Flash-Next (xhigh, pp=2) | stop | 86,377 | 88 KB | Yes |
| Regular Flash-Next | stop | 78,757 | 64 KB | Yes |
All six runs produced the full three-file deliverable (index.html, style.css, script.js) with proper HTML structure, canvas rendering, game state management, and a verification framework. None of them truncated mid-code or burned their budget on planning prose (with one exception, discussed below).
The MXFP4 dense model’s output had a subtle but fatal bug: the CSS defined .screen.visible { display: flex; } to show game screens, and the JavaScript dispatched a phase-change custom event when the game state changed – but nothing ever added the .visible class. The event handler updated text content (scores, level numbers) but never toggled screen visibility.
The result: the game loaded without errors, the canvas rendered, the state machine worked internally – but the start button was invisible because every screen had display: none. Clicking where the button should be did nothing.
The fix was two lines of JavaScript. We added screen visibility toggling to the phase-change event handler (hide all screens, show the one matching the new phase) and an initial call in init() to show the menu screen on load. The full diff is included in the artifact tarball (MANUAL_FIX.diff).
Here’s the thing: if this had been run through a harness – even a simple one that loads the page, checks for a visible start button, and reports “button not found” – the model would have had the feedback it needed to fix the bug in one retry. The bug is exactly the kind a code review or automated test catches immediately. In a one-shot test with no feedback loop, it’s fatal.
The most interesting failure was Swift1.5 Flash-Next at its default reasoning effort (xhigh). The model generated 255,555 tokens – nearly a million characters of reasoning content – without producing a single line of actual code.
The reasoning content was self-referential and repetitive: the model kept listing things it “needed to ensure” about the game state, in patterns like:
“Need ensure no use of state.ghostClearOnGhostClearOnLevel? no.
Need ensure no use of state.ghostClearOnGhostClearOnLives? no.
Need ensure no use of state.ghostClearOnGhostClearOnMiss? no.”
This continued for hundreds of thousands of tokens. The model was stuck in a loop where each “need to check X” generated more “need to check Y” thoughts, never converging on actual output, but also several dozens of lines before a repeat would occur.
The fix was presence_penalty: 2.0 – a setting the Qwen model card explicitly documents as reducing “endless repetition.” With presence_penalty at 2.0 and reasoning_effort at xhigh, the model produced 86,377 tokens of reasoning (still a lot, but bounded) followed by a complete, working game.
Qwen is a “nervous test taker” ? the model wasn’t broken. It was doing exactly what we told it to do – think as deeply as possible (xhigh) – and it thought itself into a corner. The repetition penalty we needed was literally documented on the model card. We just hadn’t read that far before running the test. When we set reasoning_effort to “low,” the model produced a clean game in 20,735 tokens with only 7,380 tokens of reasoning. It stopped overthinking because we stopped telling it to overthink.
| Setting | tok/s | Reasoning tokens | Content tokens | Result |
|---|---|---|---|---|
| low | 178 | 7,380 | 20,735 | Clean code, 117 seconds |
| medium | 137 | 135,581 | 58,178 | Clean code, 424 seconds |
| xhigh + pp=2 | 110 | 247,512 | 86,377 | Clean code, 786 seconds |
| xhigh + pp=0 | 184 | 921,676 | 0 | Reasoning loop, no code |
The Qwen model card documents reasoning_effort values of xhigh (default), medium, and low. It also notes that “in multi-turn agentic tasks, lower reasoning effort doesn’t always reduce total completion time” – which is true for chat, but for code generation the opposite is clearly the case. Lower effort meant faster completion and better output, because the model spent its token budget on code instead of self-referential reasoning.
There is a great qualitative aspect to this, so here’s the output from each of the runs for science.
Each artifact tarball includes a screenshot.png from the browser showing what the game actually looks like when loaded…
The FP8 dense model’s output had an odd bug I didn’t catch at first: buildLevel() referenced CONFIG.grid.rowColors[row], but rowColors was defined at the CONFIG level, not inside CONFIG.grid. This caused a TypeError: Cannot read properties of undefined (reading '0') but I wasn’t paying attention to the console.
The fix was changing one word: g.rowColors[row] to CONFIG.rowColors[row]. Again, this is exactly the kind of bug a harness would catch immediately – the page throws an error on load, the harness reports it, the model fixes it.
Each breakout clone has minor visual glitches, edge-case bugs, and polish issues that we didn’t catalog. The games are playable but not production-quality. Some examples:
Our general impression: the less quantized the model, the better the result. The FP8 dense model (8-bit trunk) produced the cleanest code with the fewest bugs. The MXFP4 model (4-bit experts) had the screen visibility bug. The Flash-Next models (4-bit experts + int8 trunk) produced working games but with more reasoning overhead. This is a sample size of one prompt, but I thought it was a good cross-section of my experiments on AMD hardware with Radiance.
And a good harness makes the result that much better. Every bug we found – the MXFP4 screen visibility, the fp8_dense rowColors reference, the Flash-Next reasoning loop – would have been caught and fixed in one or two retry cycles with even a minimal evaluation harness.
Every breakout run’s complete output is packaged for evaluation. Each tarball contains the extracted game files (index.html, style.css, script.js), a breakout.html you can double-click to play, the raw model output for reference, and metadata about the model, settings, and results.
[fp8_dense.tar.gz](https://forum.level1techs.com/uploads/short-url/9wGykZ70XfNR4EfJk4k9hbUjLZy.gz) (63.3 KB)
[mxfp4_dense.tar.gz](https://forum.level1techs.com/uploads/short-url/z8ozQNsP471SyWebWRaewh0CT7w.gz) (69.0 KB)
[regular_fn.tar.gz](https://forum.level1techs.com/uploads/short-url/sf156Q3fHlV9ns1pP9y3bt5J4Uo.gz) (58.6 KB)
[swift_fn_low.tar.gz](https://forum.level1techs.com/uploads/short-url/qN6x9KzVwU8TDZL67zCuLRvhKR2.gz) (62.7 KB)
[swift_fn_medium.tar.gz](https://forum.level1techs.com/uploads/short-url/nJJBb9cWkfKNYEKo39zQw7RvyIa.gz) (113.6 KB)
[swift_fn_xhigh.tar.gz](https://forum.level1techs.com/uploads/short-url/cMGazTbOQqe3qiYWYMqSHagFnEy.gz) (324.8 KB)
| Artifact | Model | Size | Notes |
|---|---|---|---|
fp8_dense.tar.gz |
Swift1.5-Qwen3.8-27B FP8 | 65 KB | Includes the rowColors fix |
mxfp4_dense.tar.gz |
Qwen3.8-27B MXFP4 | 71 KB | Includes the screen visibility fix + diff |
swift_fn_low.tar.gz |
Swift1.5 Flash-Next (low effort) | 64 KB | Cleanest run, no fixes needed |
swift_fn_medium.tar.gz |
Swift1.5 Flash-Next (medium effort) | 116 KB | Includes reasoning trace |
swift_fn_xhigh.tar.gz |
Swift1.5 Flash-Next (xhigh, pp=2) | 333 KB | Includes 248 KB reasoning trace |
regular_fn.tar.gz |
Qwen3.8-Flash-Next FP8+IQ4R | 60 KB | No fixes needed |
breakout_prompt.txt (28.9 KB)
Download the artifacts, extract, and double-click breakout.html to evaluate the games yourself. The MXFP4 tarball includes both breakout.html (the original, broken output) and breakout_with_patch.html (the fixed version) so you can see the difference.
Radiance is a standalone C++/HIP inference engine for AMD RDNA4 GPUs, written from scratch by @StillDeadcode. It started as vllm-radiance (a patched vLLM fork with hand-written RDNA4 kernels) and evolved into a fully independent engine with a plugin architecture. The Docker image is 423 MB. It supports FP8, MXFP4, and bf16 quantization, speculative decoding (MTP and DFlash2), tensor parallelism up to TP4, and expert tiering for MoE models.
For setup instructions, benchmark methodology, and performance tuning, see our other writeups elsewhere on this forum. The short version: pull stilldeadcode/radiance:latest, download a .rad model container from HuggingFace, and run with --tp 4 --tp-wire wht6 --p2p off on an 8-GPU system.
Six models, one prompt, no harness.
The throughput numbers are strong: 147-242 tok/s for single-stream code generation on consumer AMD Radeon RDNA GPUs, with aggregate throughput up to 1103 tok/s at concurrency (2x R9700s) (don’t wanna spoil the peak numbers from 8x R9600s just yet haha!). The qualitative results are encouraging: these models can build complete, runnable applications from a detailed specification in a single pass.