AM4 motherboard with quad R9700's, Proxmox VM + plx88096 P2P enablement Woes A user on the Level1Techs forum reported getting four AMD R9700 AI GPUs working behind a PLX88096 PCIe switch on an AM4 X570-D4U-2L2T/BCM motherboard running Proxmox, after enabling Legacy CSM support resolved the initial BAR space failure that had also knocked out the onboard 10GbE. Peer-to-peer (P2P) remains disabled inside the passed-through VM despite PCIe atomics showing enabled on all four cards in both host and guest, and the user plans to test an unprivileged LXC container next. The system currently runs Qwen 3.8 flash-next at roughly 30 tokens per second with each GPU drawing under 100 of its 300 watts. EDIT: So i managed to get my system working with all 4 GPU’s behind the pcie switch, however i still can’t get the damn P2P to work nice inside my VM’s, i’m suspecting it has to do with PCIE atomics not working properly for guest vm’s or something of that nature, can anyone more experienced in this than me chip-in? So i recently got my PLX88096 chip in the mail, and have used it to install 4 R9700 AI’s onto the system, however it seems my poor prosumer grade AM4 X570-D4U-2L2T/BCM motherboard has had enough of my shenanigans and just doesnt have the BAR space to deal with it My onboard 10Gbe also stopped working, but i suspect this is just a side effect of the lack of BAR space I’ve added the Dmesg logs and my kernel opts Please obi-wan janitobi’s, you guys are my only hope: Kernel Args: GRUB CMDLINE LINUX DEFAULT=“quiet amd iommu=on iommu=pt pcie aspm=off amdgpu.ras enable=0 pci=realloc=on amdgpu.runpm=0 amdgpu.gpu recovery=1 pcie acs override=downstream,multifunction, default hugepagesz=1G hugepagesz=1G hugepages=108 transparent hugepage=always” DMESG: See attachment dmesg.txt https://forum.level1techs.com/uploads/short-url/cuEyL12crKwewszQIO6ydsM2IAr.txt 46.6 KB For your viewing pleasure or distaste, depending on tolerance : 3 Likes It WORKS , it was Legacy CSM support being enabled, i have a new problem now though, even though everything is now properly showing inside the VM P2P is still disabled in AMD-SMI topology: ACCESS TABLE: 0000:01:00.0 0000:02:00.0 0000:03:00.0 0000:04:00.0 0000:01:00.0 ENABLED DISABLED DISABLED DISABLED 0000:02:00.0 DISABLED ENABLED DISABLED DISABLED 0000:03:00.0 DISABLED DISABLED ENABLED DISABLED 0000:04:00.0 DISABLED DISABLED DISABLED ENABLED WEIGHT TABLE: this is inside the VM that’s being passed through all these devices So i managed to get my system working with all 4 GPU’s behind the pcie switch passed through to my VM, however i still can’t get the damn P2P to work nice inside my VM’s, i’m suspecting it has to do with PCIE atomics not working properly for guest vm’s or something of that nature, can anyone more experienced in this than me chip-in? Alright i’m starting to suspect QEMU p2p within proxmox is a non-starter due to atomics issues, within Proxmox atomics are perfectly enabled on all four cards: lspci -s 33:00.0 -vvv | grep -i atom AtomicOpsCap: 32bit+ 64bit+ 128bitCAS- AtomicOpsCtl: ReqEn+ ECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr- ECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr- ECRC- UnsupReq- ACSViol- UncorrIntErr+ BlockedTLP- AtomicOpBlocked- TLPBlockedErr- Tree also looks good: lspci -tv: lspcitv.txt https://forum.level1techs.com/uploads/short-url/zrEUDAokYcd9VuOKdpfqbMbSwF6.txt 5.7 KB When i get home from work I will try and make an unprivileged LXC container instead to see if the Atomics magically start to work. If there is any wizened proxmox PCIe guru out there, i’d appreciate your thoughts XD, i’d prefer to use VM’s still even if LXC does work On the flip-side though i’m already running Qwen 3.8 flash-next at a reasonable 30-ish tokens a sec but none of the GPU’s are using more than 100 watts out of 300, so i’m leaving WAAYYY too much performance on tap You you have above 4g decoding and rebar enabled? Oh you have that server style am4 board, that should have adequate bar space then yeah the GPU’s actually work great now, just no P2P behind my pex88096 still i though atomics were disabled in the guest vm but they actually DO seem to be enabled on closer inspection on the guest VM: root@Rocmbuntu-V2:~ lspci -s 04:00.0 -vvv | grep -i atom AtomicOpsCap: 32bit+ 64bit+ 128bitCAS- AtomicOpsCtl: ReqEn+ ECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr- ECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr- ECRC- UnsupReq- ACSViol- UncorrIntErr+ BlockedTLP- AtomicOpBlocked- TLPBlockedErr- root@Rocmbuntu-V2:~ lspci -s 03:00.0 -vvv | grep -i atom AtomicOpsCap: 32bit+ 64bit+ 128bitCAS- AtomicOpsCtl: ReqEn+ ECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr- ECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr- ECRC- UnsupReq- ACSViol- UncorrIntErr+ BlockedTLP- AtomicOpBlocked- TLPBlockedErr- root@Rocmbuntu-V2:~ lspci -s 02:00.0 -vvv | grep -i atom AtomicOpsCap: 32bit+ 64bit+ 128bitCAS- AtomicOpsCtl: ReqEn+ ECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr- ECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr- ECRC- UnsupReq- ACSViol- UncorrIntErr+ BlockedTLP- AtomicOpBlocked- TLPBlockedErr- root@Rocmbuntu-V2:~ lspci -s 01:00.0 -vvv | grep -i atom AtomicOpsCap: 32bit+ 64bit+ 128bitCAS- AtomicOpsCtl: ReqEn+ ECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr- ECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr- ECRC- UnsupReq- ACSViol- UncorrIntErr+ BlockedTLP- AtomicOpBlocked- TLPBlockedErr- even the lspci -tv looks clean inside the VM to me: lspcitvqemuguest.txt https://forum.level1techs.com/uploads/short-url/kukLWlQy960YE65girxpJMKsu4f.txt 2.3 KB So if its not atomics or motherboard going through the pex88096 switch holding back p2p, what could it be That’s all linux speak to me I just know about bar space because I use enterprise cards to play games LOL Have you enabled sr-iov? Seems like that’s an important vm thing Kind of hoping Wendel or someone else knowledgeable on P2P enablement can give me some tests to run, i’m running out of ideas and i’m not feeling too hopeful about LXC now that i know atomics probably aint it A quick Google on amd p2p Do you have the gpus running through a PCI-E muxer card? That might interfere with the uhhh “directness” of the direct memory access Try take a card out and putting it directly into one of the slots and see if it complains on every card but one I’m behind a pcie switch specifically a PLX88096, direct CPU attachment also does not work have already tried , the switch should not be the problem though if anything it should help due to it allowing bypass of the AM4 consumer chipset platform, if you search through the forum its an often recommended remedy to use a switch like this to enable P2P on prosumer hardware…..just not for my setup unfortunately RCCL might be broken if you’re not behind a switch, but direct peer-to-peer is also iffy if you’re in a virtual machine. Digital Spaceport has a guide on running VLLM in a container in LXC and Proxmox, and I think you might want to give that a try. Just to see. Obviously the container setup is going to be pretty similar, but the software is probably going to be very different because you’re on RDNA 4 and not using CUDA. That said, if you can get it working in a VM, you probably can get it working on the host. And it’s really just because you need peer-to-peer. I’m running a virtual machine with four 9700 Pros myself. I’m just not using direct peer-to-peer because no switch, it wouldn’t work anyway. Yeah i wasn’t expecting it to work without the switch, apart from being able to run 4 of them getting P2P potentially working was a major reason for buying the PEX88096 switch from aliexpress I’m going to try and get it to work with LXC tonight, wish me luck root@Rocmbuntu-V2:~ amd-smi topology ACCESS TABLE: 0000:01:00.0 0000:01:00.1 0000:01:00.2 0000:01:00.3 0000:01:00.0 ENABLED ENABLED ENABLED ENABLED 0000:01:00.1 ENABLED ENABLED ENABLED ENABLED 0000:01:00.2 ENABLED ENABLED ENABLED ENABLED 0000:01:00.3 ENABLED ENABLED ENABLED ENABLED WEIGHT TABLE: THANK YOU DEEPSEEK V4.1 FLASH XD uh one moment now the rocm bandwidth test won’t run, performing transfers resets the GPU gpu-reset.txt https://forum.level1techs.com/uploads/short-url/zLmTLThBpqYXbvLF4CZgFpAB7Ao.txt 3.0 KB IT WORKS For the most part i’m only getting 28.5GB/s bidi between cards, not the 63~ i’d expect from pcie gen 4 , Still though i’m over the effing moon, since atleast i can try out VLLM properly now on my VM setup Common Suppress by setting HIDE ENV=1 ALWAYS VALIDATE = 0 : Validating after all iterations BLOCK BYTES = 256 : Each CU gets a mulitple of 256 bytes to copy BYTE OFFSET = 0 : Using byte offset of 0 CU MASK = 0 : All FILL COMPRESS = 0 : Not specified FILL PATTERN = 0 : Element i = i 517 modulo 383 + 31 srcBufferIdx + 1 GFX BLOCK ORDER = 0 : Thread block ordering: Sequential GFX BLOCK SIZE = 256 : Threadblock size of 256 GFX SINGLE TEAM = 1 : Combining CUs to work across entire data array GFX TEMPORAL = 0 : Not using non-temporal loads/stores GFX UNROLL = 4 : Using GFX unroll factor of 4 GFX WAVE ORDER = 0 : Using GFX wave ordering of Unroll,Wavefront,CU GFX WORD SIZE = 4 : Using GFX word size of 4 DWORDx4 MIN VAR SUBEXEC = 1 : Using at least 1 subexecutor s for variable subExec tranfers MAX VAR SUBEXEC = 0 : Using up to all available subexecutors for variable subExec transfers NUM ITERATIONS = 10 : Running 10 timed iteration s NUM SUBITERATIONS = 1 : Running 1 subiterations NUM WARMUPS = 3 : Running 3 warmup iteration s per Test SHOW ITERATIONS = 0 : Hiding per-iteration timing USE HIP EVENTS = 1 : Using HIP events for GFX/DMA Executor timing USE HSA DMA = 0 : Using hipMemcpyAsync for DMA execution USE INTERACTIVE = 0 : Running in non-interactive mode USE SINGLE STREAM = 1 : Using single stream per GFX device VALIDATE DIRECT = 0 : Validate GPU destination memory via CPU staging buffer VALIDATE SOURCE = 0 : Do not perform source validation after prep P2P Related NUM CPU DEVICES = 1 : Using 1 CPUs NUM CPU SE = 4 : Using 4 CPU threads per Transfer NUM GPU DEVICES = 4 : Using 4 GPUs NUM GPU SE = 32 : Using 32 GPU subexecutors/CUs per Transfer P2P MODE = 0 : Running Uni + Bi transfers USE FINE GRAIN = 0 : Using coarse-grained memory USE GPU DMA = 0 : Using GPU-GFX as GPU executor USE REMOTE READ = 0 : Using SRC as executor Bytes Per Direction 268435456 Unidirectional copy peak bandwidth GB/s Local read / Remote write GPU-Executor: GFX SRC+EXE\DST CPU 00 GPU 00 GPU 01 GPU 02 GPU 03 CPU 00 → 16.17 11.17 11.18 11.19 11.17 GPU 00 → 14.27 297.29 14.24 14.25 14.26 GPU 01 → 14.27 14.26 299.37 14.25 14.25 GPU 02 → 14.25 14.26 14.26 299.89 14.25 GPU 03 → 14.28 14.26 14.25 14.25 299.76 CPU- CPU CPU- GPU GPU- CPU GPU- GPU Averages During UniDir : N/A 11.18 14.27 14.25 Bidirectional copy peak bandwidth GB/s Local read / Remote write GPU-Executor: GFX SRC\DST CPU 00 GPU 00 GPU 01 GPU 02 GPU 03 CPU 00 → N/A 11.08 11.07 11.07 11.09 CPU 00 ← N/A 14.19 14.19 14.19 14.19 CPU 00 ↔ N/A 25.28 25.26 25.26 25.27 GPU 00 → 14.19 N/A 13.89 13.88 13.90 GPU 00 ← 11.08 N/A 13.88 13.88 13.89 GPU 00 ↔ 25.27 N/A 27.77 27.77 27.79 GPU 01 → 14.19 13.88 N/A 13.88 13.89 GPU 01 ← 11.08 13.90 N/A 13.90 13.88 GPU 01 ↔ 25.27 27.78 N/A 27.78 27.77 GPU 02 → 14.17 13.89 13.89 N/A 13.89 GPU 02 ← 11.08 13.87 13.87 N/A 13.87 GPU 02 ↔ 25.25 27.76 27.76 N/A 27.76 GPU 03 → 14.19 13.88 13.89 13.87 N/A GPU 03 ← 11.10 13.89 13.89 13.89 N/A GPU 03 ↔ 25.29 27.77 27.79 27.77 N/A CPU- CPU CPU- GPU GPU- CPU GPU- GPU Averages During BiDir : N/A 12.63 12.64 13.89 I cant say i did this solo thanks deepseek :P But the problem was two-fold: The Rocm DKMS module has a bug in it, that just straight up disables P2P, after deepseek corrected that the ENABLED messages started but i still could not use actual tools like rocm bandwidth test as it would just cause GPU resets. The Magic bullet turned out to be lowering the position in-memory of the GPU’s see the attached AI writeup After setting this i was finally able to use P2P tests and tooling, My next target is either figuring out why i seem to be stuck at pcie gen3 speeds and running VLLM with RCCL P2P-SETUP.txt https://forum.level1techs.com/uploads/short-url/hIAaUM8bXhlTHiuNul1lZV80DrJ.txt 5.6 KB 1 Like