Advice on a local AI server before buying A forum commenter running Qwen 3.8 Flash Next on dual AMD R9700 GPUs reported roughly 200 tokens per second decode and 6200 tokens per second prefill at 4-bit precision, and advised buyers to use an x8/x8 motherboard configuration because placing an R9700 in a chipset slot cost that card about 35-40% performance from added latency. The commenter also said llama.cpp is currently one of the slowest ways to run models given its problems with speculative decoding and concurrency. Can you stretch the budget to go for the R9700 instead, if you drop to a used Ryzen 5x00 and an X570 board that’ll do x8/x8? For what it’s worth, the radiance engine is currently running Qwen 3.8 Flash Next on my dual R9700s at ~200t/s decode, 6200t/s prefill 4-bit, with FP8 ngram table and activations, I think . I don’t know how it’d run with just one I’m using mine right now, so I can’t check , but I’d hazard a guess that it’s significantly faster than a 3090. BTW, yes…if you want to go for a second card and TP, you’re going to want x8/x8. Layer split doesn’t need that, but at the very least you’re leaving a lot of performance on the table. With R9700s and one in a chipset slot, that card lost ~35-40% performance from the extra hop of latency. Also…with the advent of all these architecture-specific inference engines, llama.cpp is currently one of the slowest ways to run models also incredibly limited, given its problems with spec decoding and concurrency .