P2P enablement on prosumer AM4 with Quad AMD R9700's....on a proxmox VM with a PLX switch A prosumer running four AMD R9700 GPUs on an AM4 platform inside a Proxmox VM with a PLX switch got peer-to-peer inference working, reaching 82GB of host RAM usage, roughly 110GB of VRAM and a 500k-token GPU KV cache. After disabling debug flags that rebuilt graphs per token, the builder reported throughput of 27.5 tokens per second single-user and 300 tokens per second concurrent, up from 4 tk/s single-user and 100 tk/s concurrent. The builder noted the R9700's FP8 matrix multiplication strength and advised avoiding W4A16 quantization in favor of formats such as tcclaviger/Qwen3.8-Flash-Next-MXFP4-FP8-GPTQ on Hugging Face. i’ve added a 100GB of swap space as a hail-mary, i hope this lets me get past the transient spikes I got it running , benches will follow, stupid dummy load tester blew up hostram causing OOM’s not seen with the actual effing model ok so i’m loaded to 82GB of host RAM and around 110GB of VRAM with a GPU kv cache of 500k, it runs but theres some weirdness: 100 tk/s in concurrent which is …..fine and might be expected considering n-gram streaming from ram, however: Single user drops to a paltry 4 tk/s, investigating further turns out i had some debug flags enabled that caused the graphs to get rebuilt per token, I’m at 27.5 tg/s single and 300 tg/s concurrent now Have you taken a look at tcclaviger/Qwen3.8-Flash-Next-MXFP4-FP8-GPTQ on huggingface? The big strength of the R9700 is FP8 matrix multiplication so you want to avoid W4A16, which is best for older Nvidia/AMD hardware.