Couldn't resist grabbing a CMP 170HX, and now I'm in a sticky position A user who purchased a CMP 170HX mining GPU for AI workloads reports that it performs comparably to a pair of Radeon R9700 GPUs in decoding but has significantly slower prefill performance, and it throttles due to thermal limits. The user is considering selling either the CMP 170HX or the R9700s, with potential recoup of £1.2k or £2k respectively, and notes difficulties with CUDA setup in llama.cpp. So…yeah, I paid more than it might be worth for a CMP 170HX. Printed a shroud, stuck a decent-ish 80mm fan on there 3000rpm, apparently optimised for static pressure , did the unlock dance for the full 64GB, and… Damn. This thing’s actually pretty close to my pair of R9700s, under specific circumstances. Basically, it’s comparable on decode performance almost identical, in fact , but the prefill lets it down - somewhere around 50% of the R9700s, in fact. …until MTP gets involved on both setups. Suddenly, the prefill is only about 20% slower than the R9700s, and prefill remains pretty much identical both with Qwen 3.6 35B and 27B Vulkan on the R9700s, CUDA on the 170HX . Now, the R9700 cooling is vastly better - this thing hits its 85C limit and throttles within a couple of minutes, and performance drops accordingly. I don’t know how much of that is the janky fan-and-shroud cooling setup, how much is the fact that it’s most likely in dire need of a repaste, and how much is just the thermal limit of the hardware. So…I’m obviously left with a bit of a quandary. I originally got it just to see if it would work, and kinda hoped I’d come up with some great use case for two AI servers at home - multiple agents etc - but I’m not sure I can, given my usual use case of mainly just coding. So what do I do? It may well come down to selling one setup or the other…but which one? I could probably recoup £2k by selling the R9700s and living with the reduced prefill and much increased load times, thanks to PCIE 2.0 x4 , or £1.2k by selling the 170HX and just going back to the original setup. What would you do? EDIT: As an aside, running llama.cpp in CUDA isn’t the easy life I thought it was - I had all sorts of pain getting it working, from broken installers direct from Nvidia to the fact that llama.cpp binary releases don’t actually include CUDA at all