cd /news/large-language-models/testing-radiance-on-2-4-8-r9700s-and… · home › topics › large-language-models › article
[ARTICLE · art-149302] src=forum.level1techs.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Testing Radiance on 2-4-8 R9700s (and R9600Ds): Some Qualitative Breakout Fun

Five of six Qwen3.8-family model runs on AMD RDNA4 GPUs with the Radiance inference engine produced a playable Breakout game from a single 6,567-token prompt with no harness or retry loop, according to a Level1Techs writeup by author StillDeadcode. The tests ran on a 2x ASUS R9700 system and an 8x ASROCK R9600D server, with a 262,144-token budget; the MXFP4 dense model's output needed a two-line fix because its JavaScript never added the `.visible` class that CSS defined for showing game screens. Throughput ranged from 250 to 324 aggregate tokens per second at concurrency 8, with the MXFP4 dense model reaching 678 tok/s peak aggregate and Flash-Next MoE hitting 1,103 tok/s at concurrency 16.

read9 min views3 publishedOct 11, 2026

This writeup covers the qualitative side of testing Qwen3.8-family models on AMD RDNA4 GPUs with the Radiance inference engine. Our earlier posts covered the performance numbers and setup instructions. This one is about what you can actually expect when you use these models – how they behave, where they break, what the output looks like, and what you need to know about reasoning settings.

If you want the how-to guides and benchmark methodology, check out our other writeups:

For testing, we used the 2x ASUS R9700 system and our 8x ASROCK R9600D server system.

( btw, errybody welcome [@StillDeadcode](https://forum.level1techs.com/u/stilldeadcode) , the author )

We used a 6,567-token programming prompt that asks the model to build a complete Breakout game in HTML/CSS/JavaScript, with a built-in 7-test verification framework. The prompt specifies either a single-file or three-file layout and includes detailed requirements for game mechanics, state management, and event-driven architecture.

This is a one-shot test. No harness (so the verification framework doesn’t run in this test – more on that in a bit), no retry loop, no feedback cycle. The model gets the prompt, generates until it hits the token budget or emits an EOS token, and whatever it produces is what we evaluate. This is deliberately harder than a harness-based evaluation – a harness would catch the “start button does nothing” class of bug and prompt the model to fix it. We wanted to see what falls out raw before passing to a harness that can run the verification criteria.

The token budget was 262,144 (the model’s full native context). Every run had enough room to finish.

Before the qualitative results, here are the raw throughput numbers for context.

Model Aggregate tok/s (conc=8) Single-stream sustained Breakout tok/s
FP8 dense (Swift1.5 27B) 324 179 tok/s 147
MXFP4 dense (Qwen3.8-27B) 322 – 242
Swift1.5 Flash-Next MoE 274 175 tok/s 178 (low effort)
Regular Flash-Next MoE 250 162 tok/s 171
Model Single-stream sustained Peak aggregate
Swift1.5 FP8 dense 124 tok/s 586 @ conc8
Qwen3.8-27B FP8 dense 130 tok/s 596 @ conc8
Qwen3.8-27B MXFP4 dense – 678 @ conc8
Flash-Next MoE 137 tok/s 1103 @ conc16

The dense models are faster in single-stream throughput because they have less total weight to stream per token. Flash-Next compensates at higher concurrency because its MoE architecture activates only 6B parameters per token, leaving more VRAM for KV cache and batch slots.

Five out of six models produced a playable game on the first shot. No harness, no retry, no code review loop. The model read a 6,567-token specification and produced a complete, runnable Breakout clone.

Model Finish reason Tokens generated Output size Game playable?
FP8 dense (Swift1.5) stop 58,416 61 KB Yes
MXFP4 dense stop 67,988 65 KB Yes (after two-line fix)
Swift1.5 Flash-Next (low effort) stop 20,735 74 KB Yes
Swift1.5 Flash-Next (medium effort) stop 58,178 80 KB Yes
Swift1.5 Flash-Next (xhigh, pp=2) stop 86,377 88 KB Yes
Regular Flash-Next stop 78,757 64 KB Yes

All six runs produced the full three-file deliverable (index.html, style.css, script.js) with proper HTML structure, canvas rendering, game state management, and a verification framework. None of them truncated mid-code or burned their budget on planning prose (with one exception, discussed below).

The MXFP4 dense model’s output had a subtle but fatal bug: the CSS defined .screen.visible { display: flex; } to show game screens, and the JavaScript dispatched a phase-change custom event when the game state changed – but nothing ever added the .visible class. The event handler updated text content (scores, level numbers) but never toggled screen visibility.

The result: the game loaded without errors, the canvas rendered, the state machine worked internally – but the start button was invisible because every screen had display: none. Clicking where the button should be did nothing.

The fix was two lines of JavaScript. We added screen visibility toggling to the phase-change event handler (hide all screens, show the one matching the new phase) and an initial call in init() to show the menu screen on load. The full diff is included in the artifact tarball (MANUAL_FIX.diff).

Here’s the thing: if this had been run through a harness – even a simple one that loads the page, checks for a visible start button, and reports “button not found” – the model would have had the feedback it needed to fix the bug in one retry. The bug is exactly the kind a code review or automated test catches immediately. In a one-shot test with no feedback loop, it’s fatal.

The most interesting failure was Swift1.5 Flash-Next at its default reasoning effort (xhigh). The model generated 255,555 tokens – nearly a million characters of reasoning content – without producing a single line of actual code.

The reasoning content was self-referential and repetitive: the model kept listing things it “needed to ensure” about the game state, in patterns like:

“Need ensure no use of state.ghostClearOnGhostClearOnLevel? no.

Need ensure no use of state.ghostClearOnGhostClearOnLives? no.

Need ensure no use of state.ghostClearOnGhostClearOnMiss? no.”

This continued for hundreds of thousands of tokens. The model was stuck in a loop where each “need to check X” generated more “need to check Y” thoughts, never converging on actual output, but also several dozens of lines before a repeat would occur.

The fix was presence_penalty: 2.0 – a setting the Qwen model card explicitly documents as reducing “endless repetition.” With presence_penalty at 2.0 and reasoning_effort at xhigh, the model produced 86,377 tokens of reasoning (still a lot, but bounded) followed by a complete, working game.

Qwen is a “nervous test taker” ? the model wasn’t broken. It was doing exactly what we told it to do – think as deeply as possible (xhigh) – and it thought itself into a corner. The repetition penalty we needed was literally documented on the model card. We just hadn’t read that far before running the test. When we set reasoning_effort to “low,” the model produced a clean game in 20,735 tokens with only 7,380 tokens of reasoning. It stopped overthinking because we stopped telling it to overthink.

Setting tok/s Reasoning tokens Content tokens Result
low 178 7,380 20,735 Clean code, 117 seconds
medium 137 135,581 58,178 Clean code, 424 seconds
xhigh + pp=2 110 247,512 86,377 Clean code, 786 seconds
xhigh + pp=0 184 921,676 0 Reasoning loop, no code

The Qwen model card documents reasoning_effort values of xhigh (default), medium, and low. It also notes that “in multi-turn agentic tasks, lower reasoning effort doesn’t always reduce total completion time” – which is true for chat, but for code generation the opposite is clearly the case. Lower effort meant faster completion and better output, because the model spent its token budget on code instead of self-referential reasoning.

There is a great qualitative aspect to this, so here’s the output from each of the runs for science.

Each artifact tarball includes a screenshot.png from the browser showing what the game actually looks like when loaded…

The FP8 dense model’s output had an odd bug I didn’t catch at first: buildLevel() referenced CONFIG.grid.rowColors[row], but rowColors was defined at the CONFIG level, not inside CONFIG.grid. This caused a TypeError: Cannot read properties of undefined (reading '0') but I wasn’t paying attention to the console.

The fix was changing one word: g.rowColors[row] to CONFIG.rowColors[row]. Again, this is exactly the kind of bug a harness would catch immediately – the page throws an error on load, the harness reports it, the model fixes it.

Each breakout clone has minor visual glitches, edge-case bugs, and polish issues that we didn’t catalog. The games are playable but not production-quality. Some examples:

Our general impression: the less quantized the model, the better the result. The FP8 dense model (8-bit trunk) produced the cleanest code with the fewest bugs. The MXFP4 model (4-bit experts) had the screen visibility bug. The Flash-Next models (4-bit experts + int8 trunk) produced working games but with more reasoning overhead. This is a sample size of one prompt, but I thought it was a good cross-section of my experiments on AMD hardware with Radiance.

And a good harness makes the result that much better. Every bug we found – the MXFP4 screen visibility, the fp8_dense rowColors reference, the Flash-Next reasoning loop – would have been caught and fixed in one or two retry cycles with even a minimal evaluation harness.

Every breakout run’s complete output is packaged for evaluation. Each tarball contains the extracted game files (index.html, style.css, script.js), a breakout.html you can double-click to play, the raw model output for reference, and metadata about the model, settings, and results.

[fp8_dense.tar.gz](https://forum.level1techs.com/uploads/short-url/9wGykZ70XfNR4EfJk4k9hbUjLZy.gz) (63.3 KB)

[mxfp4_dense.tar.gz](https://forum.level1techs.com/uploads/short-url/z8ozQNsP471SyWebWRaewh0CT7w.gz) (69.0 KB)

[regular_fn.tar.gz](https://forum.level1techs.com/uploads/short-url/sf156Q3fHlV9ns1pP9y3bt5J4Uo.gz) (58.6 KB)

[swift_fn_low.tar.gz](https://forum.level1techs.com/uploads/short-url/qN6x9KzVwU8TDZL67zCuLRvhKR2.gz) (62.7 KB)

[swift_fn_medium.tar.gz](https://forum.level1techs.com/uploads/short-url/nJJBb9cWkfKNYEKo39zQw7RvyIa.gz) (113.6 KB)

[swift_fn_xhigh.tar.gz](https://forum.level1techs.com/uploads/short-url/cMGazTbOQqe3qiYWYMqSHagFnEy.gz) (324.8 KB)
Artifact Model Size Notes
fp8_dense.tar.gz Swift1.5-Qwen3.8-27B FP8 65 KB Includes the rowColors fix
mxfp4_dense.tar.gz Qwen3.8-27B MXFP4 71 KB Includes the screen visibility fix + diff
swift_fn_low.tar.gz Swift1.5 Flash-Next (low effort) 64 KB Cleanest run, no fixes needed
swift_fn_medium.tar.gz Swift1.5 Flash-Next (medium effort) 116 KB Includes reasoning trace
swift_fn_xhigh.tar.gz Swift1.5 Flash-Next (xhigh, pp=2) 333 KB Includes 248 KB reasoning trace
regular_fn.tar.gz Qwen3.8-Flash-Next FP8+IQ4R 60 KB No fixes needed

breakout_prompt.txt (28.9 KB) Download the artifacts, extract, and double-click breakout.html to evaluate the games yourself. The MXFP4 tarball includes both breakout.html (the original, broken output) and breakout_with_patch.html (the fixed version) so you can see the difference.

Radiance is a standalone C++/HIP inference engine for AMD RDNA4 GPUs, written from scratch by @StillDeadcode. It started as vllm-radiance (a patched vLLM fork with hand-written RDNA4 kernels) and evolved into a fully independent engine with a plugin architecture. The Docker image is 423 MB. It supports FP8, MXFP4, and bf16 quantization, speculative decoding (MTP and DFlash2), tensor parallelism up to TP4, and expert tiering for MoE models.

For setup instructions, benchmark methodology, and performance tuning, see our other writeups elsewhere on this forum. The short version: pull stilldeadcode/radiance:latest, download a .rad model container from HuggingFace, and run with --tp 4 --tp-wire wht6 --p2p off on an 8-GPU system. Six models, one prompt, no harness.

The throughput numbers are strong: 147-242 tok/s for single-stream code generation on consumer AMD Radeon RDNA GPUs, with aggregate throughput up to 1103 tok/s at concurrency (2x R9700s) (don’t wanna spoil the peak numbers from 8x R9600s just yet haha!). The qualitative results are encouraging: these models can build complete, runnable applications from a detailed specification in a single pass.

── more in #large-language-models 4 stories · sorted by recency
── more on @radiance 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/testing-radiance-on-…] indexed:0 read:9min 2026-10-11 · —