AI inference system recommendations: 8x R9700 vs 2x DGX Spark A user seeking advice on building a high-end AI inference system reports that their 8x R9700 setup underperforms compared to a single DGX Spark, citing 10 tokens per second on their system versus 21.65 tok/s claimed for two DGX Spark units in a GitHub repository, and asks for guidance on hardware and software support. Hello guys, I am looking to build a high end system for AI inference. I badly need your advice since there is a lot of conflicting floating around. Also long term software support and quality is a consideration. The hardware I have these GPUs running in might also be a bottle neck. I have a old Gigabyte G291-Z20 GPU https://www.gigabyte.com/Enterprise/GPU-Server/G291-Z20-rev-100 server, that only connects the GPUs on PCIe 3.0, has no ReBAR nor above 4G decoding In theory I would expect to get more throughput on 8x R9700 than the dgx spark. But I can’t seem to find any benchmarks. I already have four R9700 but my own performance tests are worse than what people report on a single dgx spark. Maybe its because my prefill is faster as everyone seems laser focused on decode speed. Like this dude on Huggingface reports getting 10 tps token per second on such a system: And then there is this repo on Github claiming 21.65 tok/s on his two dgx spark. You advice is greatly appreciated