cd /news/ai-infrastructure/budget-pcie-5-0-16x-8-card-baseboard… Β· home β€Ί topics β€Ί ai-infrastructure β€Ί article
[ARTICLE Β· art-134059] src=forum.level1techs.com β†— pub= topic=ai-infrastructure verified=true sentiment=Β· neutral

Budget Pcie 5.0 16x, 8 Card Baseboard Build

A user who purchased 8 R9700 GPUs is asking for advice on the cheapest way to build a PCIe 5.0 x16 tensor-parallel rig for running TP8 inference on ~200GB models, proposing a single Microchip Switchtec PM50100 100-lane Gen5 switch with 8 lanes per GPU and a Ryzen host with 64GB DDR5 RAM. The user cited the GitHub project local-inference-lab/rtx6kpro topology document as inspiration, contacted Guva Systems for a quote on a true 16-lanes-per-GPU switch, and is targeting a vLLM deployment on Qwen3.8-Next-Flash or quantized GLM5.3-Flash.

read2 min views1 publishedSep 18, 2026
Budget Pcie 5.0 16x, 8 Card Baseboard Build
Image: Forum (auto-discovered)

Hey all, posing the question up front: what’s the cheapest way to get a high scaling efficiency tensor parallel rig for 8 GPUs?

So, I recently impulse bought 8x r9700 with the plan to run tensor parallel tp8 on ~200GB models. It seems like for inference performance the best implementation would be to optimize for low latency and decent bandwidth through a pcie 5.0 fabric where p2p communications never hit the cpu. I’m thinking on skimping on the host as a result (cheapest cpu+mobo that exposes a pcie 16x 5.0 slot and 64gb ram).

I’m looking at the githib project β€œlocal-inference-lab/rtx6kpro/blob/master/hardware/topology.md” for inspiration using a single Microchip Switchtec PM50100 running 8x lanes per gpu flike this:

Ryzen host (64gb ddr5 ram, cheap mobo + processor)

β”‚

PCIe 5.0 x16 slot

β”‚

x16 β†’ 2Γ— MCIO x8 card

β”‚

β–Ό

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”

β”‚ C-Payne PM50100 β”‚

β”‚ 100-lane Gen5 β”‚

β”‚ switch β”‚

β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚

x8 each over MCIO

β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚

β–Ό β–Ό β–Ό β–Ό β–Ό β–Ό β–Ό β–Ό

PCIe x16 mechanical

endpoint adapters

β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚

β–Ό β–Ό β–Ό β–Ό β–Ό β–Ό β–Ό β–Ό

R9700 Γ— 8

And housing all of the bits in a small cheap gpu cluster housing like a MM-A515-CPW and swapping the electronics.

I reached out to guva systems to see if I could get a quote for a true 16x lanes per gpu switch since thats a clear optimization option, but I’m unsure what the performance unlocks for vllm optimization would be.

I’m a bit of a greenbeard in this area though, so I could use advice on hardware choice optimizations really matter for a vllm-radiance deployment on qwen3.8-next-flash or quantized glm5.3-flash like I am currently targeting.

Just realized the title is now misleading to the post content and I can’t edit it, I wandered a bit while researching this post.

── more in #ai-infrastructure 4 stories Β· sorted by recency
── more on @r9700 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/budget-pcie-5-0-16x-…] indexed:0 read:2min 2026-09-18 Β· β€”