Hello guys, I am looking to build a high end system for AI inference.
I badly need your advice since there is a lot of conflicting floating around.
Also long term software support and quality is a consideration. The hardware I have these GPUs running in might also be a bottle neck. I have a old Gigabyte G291-Z20 GPU server, that only connects the GPUs on PCIe 3.0, has no ReBAR nor above 4G decoding
In theory I would expect to get more throughput on 8x R9700 than the dgx spark. But I can’t seem to find any benchmarks.
I already have four R9700 but my own performance tests are worse than what people report on a single dgx spark. Maybe its because my prefill is faster as everyone seems laser focused on decode speed.
Like this dude on Huggingface reports getting 10 tps (token per second) on such a system:
And then there is this repo on Github claiming 21.65 tok/s on his two dgx spark.
You advice is greatly appreciated!