Can you stretch the budget to go for the R9700 instead, if you drop to a used Ryzen 5x00 and an X570 board that’ll do x8/x8?
For what it’s worth, the radiance engine is currently running Qwen 3.8 Flash Next on my dual R9700s at ~200t/s decode, 6200t/s prefill (4-bit, with FP8 ngram table and activations, I think). I don’t know how it’d run with just one (I’m using mine right now, so I can’t check), but I’d hazard a guess that it’s significantly faster than a 3090. BTW, yes…if you want to go for a second card and TP, you’re going to want x8/x8. Layer split doesn’t need that, but at the very least you’re leaving a lot of performance on the table. With R9700s and one in a chipset slot, that card lost ~35-40% performance from the extra hop of latency.
Also…with the advent of all these architecture-specific inference engines, llama.cpp is currently one of the slowest ways to run models (also incredibly limited, given its problems with spec decoding and concurrency).