{"slug": "am4-motherboard-with-quad-r9700-s-proxmox-vm-plx88096-p2p-enablement-woes", "title": "AM4 motherboard with quad R9700's, Proxmox VM + plx88096 P2P enablement Woes", "summary": "A user on the Level1Techs forum reported getting four AMD R9700 AI GPUs working behind a PLX88096 PCIe switch on an AM4 X570-D4U-2L2T/BCM motherboard running Proxmox, after enabling Legacy CSM support resolved the initial BAR space failure that had also knocked out the onboard 10GbE. Peer-to-peer (P2P) remains disabled inside the passed-through VM despite PCIe atomics showing enabled on all four cards in both host and guest, and the user plans to test an unprivileged LXC container next. The system currently runs Qwen 3.8 flash-next at roughly 30 tokens per second with each GPU drawing under 100 of its 300 watts.", "body_md": "EDIT:\n\nSo i managed to get my system working with all 4 GPU’s behind the pcie switch, however i still can’t get the damn P2P to work nice inside my VM’s, i’m suspecting it has to do with PCIE atomics not working properly for guest vm’s or something of that nature, can anyone more experienced in this than me chip-in?\n\nSo i recently got my PLX88096 chip in the mail, and have used it to install 4 R9700 AI’s onto the system, however it seems my poor prosumer grade AM4 X570-D4U-2L2T/BCM motherboard has had enough of my shenanigans and just doesnt have the BAR space to deal with it\n\nMy onboard 10Gbe also stopped working, but i suspect this is just a side effect of the lack of BAR space\n\nI’ve added the Dmesg logs and my kernel opts\n\nPlease  obi-wan janitobi’s, you guys are my only hope:\n\nKernel Args:\n\nGRUB_CMDLINE_LINUX_DEFAULT=“quiet amd_iommu=on iommu=pt pcie_aspm=off amdgpu.ras_enable=0 pci=realloc=on amdgpu.runpm=0 amdgpu.gpu_recovery=1 pcie_acs_override=downstream,multifunction, default_hugepagesz=1G hugepagesz=1G hugepages=108 transparent_hugepage=always”\n\nDMESG: See attachment\n\n[dmesg.txt](https://forum.level1techs.com/uploads/short-url/cuEyL12crKwewszQIO6ydsM2IAr.txt) (46.6 KB)\n\nFor your viewing pleasure(or distaste, depending on tolerance):\n\n \n3 Likes\n\n \nIt WORKS!!!, it was Legacy CSM support being enabled, i have a new problem now though, even though everything is now properly showing inside the VM P2P is still disabled in AMD-SMI topology:\n\nACCESS TABLE:\n\n0000:01:00.0 0000:02:00.0 0000:03:00.0 0000:04:00.0\n\n0000:01:00.0 ENABLED      DISABLED     DISABLED     DISABLED\n\n0000:02:00.0 DISABLED     ENABLED      DISABLED     DISABLED\n\n0000:03:00.0 DISABLED     DISABLED     ENABLED      DISABLED\n\n0000:04:00.0 DISABLED     DISABLED     DISABLED     ENABLED\n\nWEIGHT TABLE:\n\nthis is inside the VM that’s being passed through all these devices\n\n \n\n \nSo i managed to get my system working with all 4 GPU’s behind the pcie switch passed through to my VM, however i still can’t get the damn P2P to work nice inside my VM’s, i’m suspecting it has to do with PCIE atomics not working properly for guest vm’s or something of that nature, can anyone more experienced in this than me chip-in?\n\n \n\n \nAlright i’m starting to suspect QEMU p2p within proxmox is a non-starter due to atomics issues, within Proxmox atomics are perfectly enabled on all four cards:\n\nlspci -s 33:00.0 -vvv | grep -i atom\n\nAtomicOpsCap: 32bit+ 64bit+ 128bitCAS-\n\nAtomicOpsCtl: ReqEn+\n\nECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr-\n\nECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr-\n\nECRC- UnsupReq- ACSViol- UncorrIntErr+ BlockedTLP- AtomicOpBlocked- TLPBlockedErr-\n\nTree also looks good:\n\nlspci -tv:\n\n[lspcitv.txt](https://forum.level1techs.com/uploads/short-url/zrEUDAokYcd9VuOKdpfqbMbSwF6.txt) (5.7 KB)\n\n \n\n \nWhen i get home from work I will try and make an unprivileged LXC container instead to see if the Atomics magically start to work.\n\nIf there is any wizened proxmox PCIe guru out there, i’d appreciate your thoughts XD, i’d prefer to use VM’s still even if LXC does work\n\nOn the flip-side though i’m already running Qwen 3.8 flash-next at a reasonable 30-ish tokens a sec but none of the GPU’s are using more than 100 watts out of 300, so i’m leaving WAAYYY too much performance on tap\n\n \n\n \nYou you have above 4g decoding and rebar enabled?\n\n \n\n \nOh you have that server style am4 board, that should have adequate bar space then\n\n \n\n \nyeah the GPU’s actually work great now, just no P2P behind my pex88096 still \n\n \n\n \ni though atomics were disabled in the guest vm but they actually DO seem to be enabled on closer inspection on the guest VM:\n\nroot@Rocmbuntu-V2:~# lspci -s 04:00.0 -vvv | grep -i atom\n\nAtomicOpsCap: 32bit+ 64bit+ 128bitCAS-\n\nAtomicOpsCtl: ReqEn+\n\nECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr-\n\nECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr-\n\nECRC- UnsupReq- ACSViol- UncorrIntErr+ BlockedTLP- AtomicOpBlocked- TLPBlockedErr-\n\nroot@Rocmbuntu-V2:~# lspci -s 03:00.0 -vvv | grep -i atom\n\nAtomicOpsCap: 32bit+ 64bit+ 128bitCAS-\n\nAtomicOpsCtl: ReqEn+\n\nECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr-\n\nECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr-\n\nECRC- UnsupReq- ACSViol- UncorrIntErr+ BlockedTLP- AtomicOpBlocked- TLPBlockedErr-\n\nroot@Rocmbuntu-V2:~# lspci -s 02:00.0 -vvv | grep -i atom\n\nAtomicOpsCap: 32bit+ 64bit+ 128bitCAS-\n\nAtomicOpsCtl: ReqEn+\n\nECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr-\n\nECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr-\n\nECRC- UnsupReq- ACSViol- UncorrIntErr+ BlockedTLP- AtomicOpBlocked- TLPBlockedErr-\n\nroot@Rocmbuntu-V2:~# lspci -s 01:00.0 -vvv | grep -i atom\n\nAtomicOpsCap: 32bit+ 64bit+ 128bitCAS-\n\nAtomicOpsCtl: ReqEn+\n\nECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr-\n\nECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr-\n\nECRC- UnsupReq- ACSViol- UncorrIntErr+ BlockedTLP- AtomicOpBlocked- TLPBlockedErr-\n\neven the lspci -tv looks clean inside the VM to me:\n\n[lspcitvqemuguest.txt](https://forum.level1techs.com/uploads/short-url/kukLWlQy960YE65girxpJMKsu4f.txt) (2.3 KB)\n\nSo if its not atomics or motherboard(going through the pex88096 switch) holding back p2p, what could it be\n\n \n\n \nThat’s all linux speak to me I just know about bar space because I use enterprise cards to play games LOL\n\nHave you enabled sr-iov?\n\nSeems like that’s an important vm thing\n\n \n\n \nKind of hoping Wendel or someone else knowledgeable on P2P enablement can give me some tests to run, i’m running out of ideas and i’m not feeling too hopeful about LXC now that i know atomics probably aint it\n\n \n\n \nA quick Google on amd p2p\n\nDo you have the gpus running through a PCI-E muxer card?\n\nThat might interfere with the uhhh “directness” of the direct memory access\n\nTry take a card out and putting it directly into one of the slots and see if it complains on every card but one\n\n \n\n \nI’m behind a pcie switch specifically a PLX88096, direct CPU attachment also does not work (have already tried), the switch should not be the problem though if anything it should help due to it allowing bypass of the AM4 consumer chipset platform, if you search through the forum its an often recommended remedy to use a switch like this to enable P2P on prosumer hardware…..just not for my setup unfortunately\n\n \n\n \nRCCL might be broken if you’re not behind a switch, but direct peer-to-peer is also iffy if you’re in a virtual machine. Digital Spaceport has a guide on running VLLM in a container in LXC and Proxmox, and I think you might want to give that a try. Just to see. Obviously the container setup is going to be pretty similar, but the software is probably going to be very different because you’re on RDNA 4 and not using CUDA. That said, if you can get it working in a VM, you probably can get it working on the host. And it’s really just because you need peer-to-peer. I’m running a virtual machine with four 9700 Pros myself. I’m just not using direct peer-to-peer because no switch, it wouldn’t work anyway.\n\n \n\n \nYeah i wasn’t expecting it to work without the switch, apart from being able to run 4 of them getting P2P potentially working was a major reason for buying the PEX88096 switch from aliexpress\n\nI’m going to try and get it to work with LXC tonight, wish me luck\n\n \n\n \nroot@Rocmbuntu-V2:~# amd-smi topology\n\nACCESS TABLE:\n\n0000:01:00.0 0000:01:00.1 0000:01:00.2 0000:01:00.3\n\n0000:01:00.0 ENABLED      ENABLED      ENABLED      ENABLED\n\n0000:01:00.1 ENABLED      ENABLED      ENABLED      ENABLED\n\n0000:01:00.2 ENABLED      ENABLED      ENABLED      ENABLED\n\n0000:01:00.3 ENABLED      ENABLED      ENABLED      ENABLED\n\nWEIGHT TABLE:\n\nTHANK YOU DEEPSEEK V4.1 FLASH XD\n\n \n\n \nuh one moment now the rocm bandwidth test won’t run, performing transfers resets the GPU\n\n[gpu-reset.txt](https://forum.level1techs.com/uploads/short-url/zLmTLThBpqYXbvLF4CZgFpAB7Ao.txt) (3.0 KB)\n\n \n\n \nIT WORKS!\n\nFor the most part (i’m only getting 28.5GB/s bidi between cards, not the 63~ i’d expect from pcie gen 4),\n\nStill though i’m over the effing moon, since atleast i can try out VLLM properly now on my VM setup\n\n# \n\n[Common]                              (Suppress by setting HIDE_ENV=1)\n\nALWAYS_VALIDATE      =            0 : Validating after all iterations\n\nBLOCK_BYTES          =          256 : Each CU gets a mulitple of 256 bytes to copy\n\nBYTE_OFFSET          =            0 : Using byte offset of 0\n\nCU_MASK              =            0 : All\n\nFILL_COMPRESS        =            0 : Not specified\n\nFILL_PATTERN         =            0 : Element i = ((i * 517) modulo 383 + 31) * (srcBufferIdx + 1)\n\nGFX_BLOCK_ORDER      =            0 : Thread block ordering: Sequential\n\nGFX_BLOCK_SIZE       =          256 : Threadblock size of 256\n\nGFX_SINGLE_TEAM      =            1 : Combining CUs to work across entire data array\n\nGFX_TEMPORAL         =            0 : Not using non-temporal loads/stores\n\nGFX_UNROLL           =            4 : Using GFX unroll factor of 4\n\nGFX_WAVE_ORDER       =            0 : Using GFX wave ordering of Unroll,Wavefront,CU\n\nGFX_WORD_SIZE        =            4 : Using GFX word size of 4 (DWORDx4)\n\nMIN_VAR_SUBEXEC      =            1 : Using at least 1 subexecutor(s) for variable subExec tranfers\n\nMAX_VAR_SUBEXEC      =            0 : Using up to all available subexecutors for variable subExec transfers\n\nNUM_ITERATIONS       =           10 : Running 10  timed iteration(s)\n\nNUM_SUBITERATIONS    =            1 : Running 1 subiterations\n\nNUM_WARMUPS          =            3 : Running 3 warmup iteration(s) per Test\n\nSHOW_ITERATIONS      =            0 : Hiding per-iteration timing\n\nUSE_HIP_EVENTS       =            1 : Using HIP events for GFX/DMA Executor timing\n\nUSE_HSA_DMA          =            0 : Using hipMemcpyAsync for DMA execution\n\nUSE_INTERACTIVE      =            0 : Running in non-interactive mode\n\nUSE_SINGLE_STREAM    =            1 : Using single stream per GFX device\n\nVALIDATE_DIRECT      =            0 : Validate GPU destination memory via CPU staging buffer\n\nVALIDATE_SOURCE      =            0 : Do not perform source validation after prep\n\nP2P Related\n\nNUM_CPU_DEVICES      =            1 : Using 1 CPUs\n\nNUM_CPU_SE           =            4 : Using 4 CPU threads per Transfer\n\nNUM_GPU_DEVICES      =            4 : Using 4 GPUs\n\nNUM_GPU_SE           =           32 : Using 32 GPU subexecutors/CUs per Transfer\n\nP2P_MODE             =            0 : Running Uni + Bi transfers\n\nUSE_FINE_GRAIN       =            0 : Using coarse-grained memory\n\nUSE_GPU_DMA          =            0 : Using GPU-GFX as GPU executor\n\nUSE_REMOTE_READ      =            0 : Using SRC as executor\n\nBytes Per Direction 268435456\n\nUnidirectional copy peak bandwidth GB/s [Local read / Remote write] (GPU-Executor: GFX)\n\nSRC+EXE\\DST    CPU 00       GPU 00    GPU 01    GPU 02    GPU 03\n\nCPU 00  →     16.17        11.17     11.18     11.19     11.17\n\nGPU 00  →     14.27       297.29     14.24     14.25     14.26\n\nGPU 01  →     14.27        14.26    299.37     14.25     14.25\n\nGPU 02  →     14.25        14.26     14.26    299.89     14.25\n\nGPU 03  →     14.28        14.26     14.25     14.25    299.76\n\nCPU->CPU  CPU->GPU  GPU->CPU  GPU->GPU\n\nAverages (During UniDir):       N/A     11.18     14.27     14.25\n\nBidirectional copy peak bandwidth GB/s [Local read / Remote write] (GPU-Executor: GFX)\n\nSRC\\DST    CPU 00       GPU 00    GPU 01    GPU 02    GPU 03\n\nCPU 00  →       N/A        11.08     11.07     11.07     11.09\n\nCPU 00 ←        N/A        14.19     14.19     14.19     14.19\n\nCPU 00 ↔       N/A        25.28     25.26     25.26     25.27\n\nGPU 00  →     14.19          N/A     13.89     13.88     13.90\n\nGPU 00 ←      11.08          N/A     13.88     13.88     13.89\n\nGPU 00 ↔     25.27          N/A     27.77     27.77     27.79\n\nGPU 01  →     14.19        13.88       N/A     13.88     13.89\n\nGPU 01 ←      11.08        13.90       N/A     13.90     13.88\n\nGPU 01 ↔     25.27        27.78       N/A     27.78     27.77\n\nGPU 02  →     14.17        13.89     13.89       N/A     13.89\n\nGPU 02 ←      11.08        13.87     13.87       N/A     13.87\n\nGPU 02 ↔     25.25        27.76     27.76       N/A     27.76\n\nGPU 03  →     14.19        13.88     13.89     13.87       N/A\n\nGPU 03 ←      11.10        13.89     13.89     13.89       N/A\n\nGPU 03 ↔     25.29        27.77     27.79     27.77       N/A\n\nCPU->CPU  CPU->GPU  GPU->CPU  GPU->GPU\n\nAverages (During  BiDir):       N/A     12.63     12.64     13.89\n\n \n\n \nI cant say i did this solo (thanks deepseek :P) But the problem was two-fold:\n\nThe Rocm DKMS module has a bug in it, that just straight up disables P2P, after deepseek corrected that the ENABLED messages started but i still could not use actual tools like rocm bandwidth test as it would just cause GPU resets.\n\nThe Magic bullet turned out to be lowering the position in-memory of the GPU’s(see the attached AI writeup)\n\nAfter setting this i was finally able to use P2P tests and tooling, My next target is either figuring out why i seem to be stuck at pcie gen3 speeds and running VLLM with RCCL\n\n[P2P-SETUP.txt](https://forum.level1techs.com/uploads/short-url/hIAaUM8bXhlTHiuNul1lZV80DrJ.txt) (5.6 KB)\n\n \n1 Like", "url": "https://wpnews.pro/news/am4-motherboard-with-quad-r9700-s-proxmox-vm-plx88096-p2p-enablement-woes", "canonical_source": "https://forum.level1techs.com/t/am4-motherboard-with-quad-r9700s-proxmox-vm-plx88096-p2p-enablement-woes/257456#post_20", "published_at": "2026-10-01 17:39:38+00:00", "updated_at": "2026-10-01 18:15:07.133132+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-chips", "large-language-models"], "entities": ["AMD", "R9700", "PLX88096", "Proxmox", "AM4 X570-D4U-2L2T/BCM", "Qwen 3.8 flash-next", "QEMU", "Level1Techs"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/am4-motherboard-with-quad-r9700-s-proxmox-vm-plx88096-p2p-enablement-woes", "markdown": "https://wpnews.pro/news/am4-motherboard-with-quad-r9700-s-proxmox-vm-plx88096-p2p-enablement-woes.md", "text": "https://wpnews.pro/news/am4-motherboard-with-quad-r9700-s-proxmox-vm-plx88096-p2p-enablement-woes.txt", "jsonld": "https://wpnews.pro/news/am4-motherboard-with-quad-r9700-s-proxmox-vm-plx88096-p2p-enablement-woes.jsonld"}}