cd /news/ai-infrastructure/ultimate-2x-r9700-guide-for-qwen-3-8… · home › topics › ai-infrastructure › article
[ARTICLE · art-149093] src=forum.level1techs.com ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

Ultimate 2x R9700 guide for Qwen 3.8 27b MXFP4: ~200t/s c=1; 500t/s at c=8

A step-by-step guide details running Qwen3.8-27B dense in MXFP4 with speculative decoding on a two-card AMD Radeon AI PRO R9700 host using Radiance, a C++/HIP inference engine by StillDeadcode, achieving roughly 200 tokens per second at concurrency 1 and 500 tokens per second at concurrency 8. Radiance replaces the earlier vllm-radiance vLLM fork with a modular plugin architecture whose Docker image is 423 MB versus vLLM's 13+ GB, and the guide reports the same model running 17-23 tokens per second single-stream on llama.cpp or stock FP8. The setup requires two 32 GB gfx1201/RDNA4 cards, a Linux host with the amdgpu kernel driver, Docker or podman, and about 60 GB of free disk.

read8 min views1 publishedOct 11, 2026
Ultimate 2x R9700 guide for Qwen 3.8 27b MXFP4: ~200t/s c=1; 500t/s at c=8
Image: Forum (auto-discovered)

#

The inference engine used here has an interesting lineage. The original project was vllm-radiance (StillDeadcode/vllm-radiance) – a patched fork of vLLM that added hand-written RDNA4 HIP kernels (libr4d) for attention, gated-delta-net, all-reduce, MXFP4, and DFlash speculative decoding. It worked well but seems like it was ultimately limited by being bolted onto vLLM’s Python/PyTorch stack. And there might have been some friction in the vLLM pull requests… haha.

The author then ditched vLLM entirely and rewrote everything from scratch as Radiance (StillDeadcode/radiance) – a modular C++/HIP inference engine with a plugin architecture. Model architectures, kernel libraries, and quantizers are all .so plugins loaded at runtime. The Docker image is just 423 MB (vs vLLM’s 13+ GB). The shared GPU kernel library libr4d which you might recognize from the vllm-radiance days (and is now bundled inside the new stand-alone Radiance).

The result is a lean, purpose-built engine for RDNA4 that delivers substantially better performance than any vLLM-based approach on these cards.

This is a step-by-step for a two-card AMD Radeon AI PRO R9700 host. By the end you will have a local OpenAI-compatible server running Qwen3.8-27B dense in MXFP4 with speculative decoding, and you will have run a real end-to-end coding benchmark (a Breakout clone from a 4,500-word spec) to prove it works. Everything here was run twice: once on my main box and once on a second, freshly set up 2x R9700 machine to make sure the steps actually replicate.

The short version of why this stack is fast: the R9700 has more compute than it can feed from its 32 GB of GDDR6, so the whole game is shrinking the bytes that move per token and getting more tokens out of each pass. MXFP4 quantization halves the weight bytes, a DFlash2 drafter gets ~3-4 accepted tokens per forward pass, and the radiance kernel stack keeps everything on the card’s own memory bus instead of crossing PCIe. Same model on llama.cpp or stock FP8 runs 17-23 tok/s single-stream; this config runs 150-200-280.

#

  • 2x AMD Radeon AI PRO R9700 (32 GB, gfx1201/RDNA4). One card works too (TP1, smaller context); the repo supports 1/2/4/8foreshadowing .
  • A Linux host with the amdgpu kernel driver working (you can see /dev/kfd and /dev/dri). I validated on CachyOS (kernel 7.1.5) and Ubuntu 24.04. You do NOT install ROCm on the host; the ROCm userspace ships inside the container image.
  • Docker (or podman). Any recent version.
  • ~60 GB free disk: 19 GB source checkpoint + 19 GB built checkpoint + 2 GB drafter + ~10 GB image. The 19 GB source is deletable after setup and the script prints the command.
  • No host Python, no HuggingFace CLI, no build tools. Setup runs everything inside the image.
  • ROCm vs Vulkan? ROCm here. It’s getting super legit on RDNA.

#

ls -la /dev/kfd /dev/dri
groups   # you want render and video in the list

If /dev/kfd is missing, the amdgpu driver is not loaded; fix that first, this guide cannot help

you there. If you are not in the render/video groups, add yourself and re-login.

#

docker pull stilldeadcode/radiance:latest

#

Radiance uses .rad container files – single-file model packages with weights, config, and metadata baked in. Pre-built containers are on Hugging Face under StillDeadcode.

curl -L -o qwen3.8-27b-fp8.rad \
  https://huggingface.co/StillDeadcode/qwen3.8-27b-fp8/resolve/main/qwen3.8-27b-fp8.rad

curl -L -o qwen3.8-27b-mxfp4.rad \
  https://huggingface.co/StillDeadcode/qwen3.8-27b-mxfp4/resolve/main/qwen3.8-27b-mxfp4.rad

curl -L -o qwen3.8-next-flash-fp8-iq4r-moe.rad \
  https://huggingface.co/StillDeadcode/qwen3.8-next-flash-fp8-iq4r-moe/resolve/main/qwen3.8-next-flash-fp8-iq4r-moe.rad

#

RENDER_GID=$(getent group render | cut -d: -f3)
VIDEO_GID=$(getent group video | cut -d: -f3)

docker run -d \
  --device /dev/kfd --device /dev/dri \
  --group-add ${RENDER_GID} --group-add ${VIDEO_GID} \
  --ipc host --network host --shm-size 32g \
  --security-opt seccomp=unconfined --cap-add=SYS_PTRACE \
  --ulimit memlock=-1:-1 \
  -e HIP_VISIBLE_DEVICES=0,1 \
  -v /path/to/qwen3.8-27b-fp8.rad:/model.rad:ro \
  --name radiance \
  stilldeadcode/radiance:latest \
  --model /model.rad \
  --tp 2 --tp-wire wht6 \
  --max-model-len 262144 \
  --max-num-seqs 8 \
  --kv-cache-dtype fp8 \
  --gpu-headroom-mib 96 \
  --num-speculative-tokens 7 \
  --host 0.0.0.0 --port 8300

RENDER_GID=$(getent group render | cut -d: -f3)
VIDEO_GID=$(getent group video | cut -d: -f3)

docker run -d \
  --device /dev/kfd --device /dev/dri \
  --group-add ${RENDER_GID} --group-add ${VIDEO_GID} \
  --ipc host --network host --shm-size 32g \
  --security-opt seccomp=unconfined --cap-add=SYS_PTRACE \
  --ulimit memlock=-1:-1 \
  -e HIP_VISIBLE_DEVICES=0,1 \
  -v /path/to/qwen3.8-next-flash-fp8-iq4r-moe.rad:/model.rad:ro \
  -v /home/w/.cache/huggingface:/root/.cache/huggingface \
  --name radiance-fn \
  stilldeadcode/radiance:latest \
  --model /model.rad \
  --tp 2 --tp-wire wht6 \
  --max-model-len 200000 \
  --max-num-seqs 8 \
  --kv-cache-dtype fp8 \
  --placement expert_tiered \
  --host-pool-mib 12288 \
  --gpu-headroom-mib 96 \
  --ngram-placement ram \
  --num-speculative-tokens 3 \
  --max-num-batched-tokens 2048 \
  --host 0.0.0.0 --port 8300

Note on first start: The first load reads the entire .rad file into VRAM and pinned host memory. For Flash-Next this is ~66 GiB of data at ~0.3 GB/s over NFS, taking about 4 minutes. Subsequent starts with the file cached by the OS are faster.

#

docker ps
curl http://localhost:8300/v1/models

#

python3 -c "
import json, time, urllib.request
t0 = time.time()
req = urllib.request.Request(
  'http://localhost:8300/v1/chat/completions',
  data=json.dumps({
    'model': 'qwen35',
    'messages': [{'role': 'user', 'content': 'Write a short story about a sentient dot matrix printer.'}],
    'max_tokens': 512,
    'temperature': 0
  }).encode(),
  headers={'Content-Type': 'application/json'},
  method='POST'
)
resp = urllib.request.urlopen(req, timeout=300)
t1 = time.time()
result = json.loads(resp.read())
tokens = result['usage']['completion_tokens']
print(f'{tokens} tokens in {t1-t0:.1f}s = {tokens/(t1-t0):.1f} tok/s')
"

#

All benchmarks use the OpenAI-compatible API with temperature 0. “Sustained decode” numbers use long generations (4096+ tokens) to amortize prefill overhead.

| Test | Result | | Single stream, 512 gen | 111 tok/s (includes prefill) | | Sustained decode, 8192 gen | 130 tok/s | | Decode-only (1-token prompt) | 169 tok/s | | conc=2 | 177 tok/s agg | | conc=4 | 291 tok/s agg | | conc=8 | 596 tok/s agg | | 2K prompt decode | 23 tok/s | | 200K prompt decode | 1.7 tok/s | | VRAM used | 15.0 GiB/GPU |

| Test | Result | | Single stream, 512 gen | 133 tok/s | | conc=4 | 448 tok/s agg | | conc=8 | 678 tok/s agg | | VRAM used | 9.7 GiB/GPU |

| Test | Result | | Single stream, 512 gen | 125 tok/s | | conc=4 | 513 tok/s agg | | conc=8 | 879 tok/s agg | | conc=16 | 1103 tok/s agg | | 20K prompt decode | 25 tok/s | | VRAM used | 21.9 GiB/GPU + 47.7 GiB host RAM (n-gram) |

In some very specific benchmaxxed configs its possible to break out (hah, get it) past 200 t/s but for real-world usage these number sare much more usable.

Over time, I also found the FP8 dense version better than MXFP4 dense.

The breakout test is a 5,899-token programming task asking the model to build a complete Breakout game in a single HTML file.

| Model | tok/s | Wall time | Result | | FP8 dense (this guide) | 125.7 | 260.8s | Complete, 32,768 tokens, hit max_tokens still generating | | MXFP4 dense (old vllm-radiance) | 191.8 | 106.6s | Complete, 20,443 tokens, early EOS, other quirks | | Flash-Next FP8 (old vllm-radiance TP8) | 42.4 | N/A | Failed – burned budget on planning |

The FP8 dense result is strong: the model stayed engaged for the full 32K budget, meaning it was productively generating code the whole time.

Flash-next FP8 can be good, but it’s much slower, spends much more time reasoning. I just included it for comparison. It is a much larger model, and will be shown in the 8xR9600D video.

#

TODO upload prompt zip here

#

The --max-model-len and --max-num-seqs flags trade off between deep context and concurrent users. The KV cache budget is roughly card_total - headroom - weights - activation_arena. For FP8 dense (15 GiB weights) at 262K context with 8 sequences, the pool backs about 37% of worst case. Reduce --max-model-len or --max-num-seqs to increase the backing ratio.

The 51B-parameter n-gram/PLE table is 47.68 GiB. Options:

  • --ngram-placement ram (recommended): Load into pinned host RAM. Requires ~48 GiB free DRAM. Fastest decode.
  • --ngram-placement disk : Stream from disk via io_uring. Uses no RAM but adds latency.
  • --ngram-placement vram : Put on GPU. Needs free VRAM (unlikely with 32 GB cards).

--tp-wire wht6 uses Walsh-Hadamard-rotated 6-bit quantization for cross-GPU communication. Lossy for messages >128 KiB (prefill and busy decode). Single-sequence decode is exact. Use --tp-wire exact for bit-exact reproducibility at ~5-10% throughput cost.

#

| Symptom | Likely cause | Fix | | E unknown option | Typo in flag name in the radiance docs | Check docker run --rm stilldeadcode/radiance:latest --help | | libavx: cannot make a ring | io_uring needs more locked memory | Add --ulimit memlock=-1:-1 to docker run | | Container exits immediately | Missing --ulimit memlock or wrong flag | Check logs with docker logs <container> | | 400 invalid_request_error | Prompt exceeds --max-model-len | Shorten prompt or increase --max-model-len |

#

Footnote for the curious: the same guide validates on the ASRock 8x R9600D box (TP2 on 2 of the 8 cards, 4 sets, with model router) produced identical KV sizing and byte-identical benchmark output, within normal run-to-run variance on speed. Two cards are the sweet spot for this model; the rest of a bigger box is better spent on other instances, other models, or scaling for many MANY users. Video on that soon.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @radiance 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ultimate-2x-r9700-gu…] indexed:0 read:8min 2026-10-11 · —