{"slug": "quad-r9700-s-on-am4-shenanigans", "title": "Quad R9700's on AM4 Shenanigans", "summary": "A user on the Level1Techs forum reported running four AMD Radeon R9700 GPUs on an AM4 platform to requantize a DeepSeek model to MXFP4 and run it in the vLLM engine, with the first layer of a 48-layer model taking 9m55s and an estimated 7h56m total to finish around 00:38:46. The user said the requantized model's performance was near identical to the original and that vLLM with prompt caching outperformed llamaserver in speed and prompt handling. The user also noted GPU-to-GPU peer-to-peer transfer was still routing through the root complex rather than the PCIe switch, which they plan to address next.", "body_md": "havent asked it about tianenmen square or the uyghurs yet though \n\n \n\n \nhonestly my company pays for deepseek and glm 5-3 flash models through API and this feels right up there, even the token rate is similar\n\n \n\n \nHate to say it, but the speed and performance difference between llamaserver and vLLM is kind of nuts, its way more intelligent with prompt caching etc, i’ve barely had to wait a second even on large prompts\n\n \n\n \nOk you know what? stretch goal:\n\nLets see if i can make an mxfp4 quant from the orcarouter version and mash it into the clavviger engine: [https://www.youtube.com/watch?v=hFDcoX7s6rE](https://www.youtube.com/watch?v=hFDcoX7s6rE)\n\n \n\n \nOK well its doing ….something, i have a feeling this will take a while\n\n \n\n \n[AA] first layer took 9m55s; estimating 7h56m for 48 layers → finishes ~00:38:46\n\n[AA] layer 0: 1539 modules (1516 with stats), 2 stage(s), capture 100.6s recapture 85.8s quant 234.8s fwd#2 74.9s io-wait 0.0s, E-moved 0, micro-batch 4\n\n[AA] PLE table: 128 mmap’d shards x 2500012 rows (torch.bfloat16), scale no, gathered on CPU\n\nSee you gents tommorow \n\n \n\n \nStretchgoal accomplished XD, requanting isnt as hard as i though performance is near identical but now i can ask it about the uyghurs\n\n \n\n \nI’m going to make a version of the qwen 3.8 27B next on mxfp4 using this engine too, i feel like i can finally use these cards to their true potential all of a sudden XD\n\n \n\n \nI think this means i did it good \n\n \n\n \nIf you install the ROCm Validation Suite, you can use the TransferBench tool to test the p2p speed GPU to GPU and GPU to CPU.\n\nI didn’t  have any issues getting p2p to work with Proxmox, but the speed was a lot lower than when running the same tests bare-metal.\n\n \n\n \nAh i know transferbench, thats how i know its still root-complex routing rather than the PCIE switch, now that VLLM and MXFP4 are sorted its my next target to fix \n\n \n1 Like", "url": "https://wpnews.pro/news/quad-r9700-s-on-am4-shenanigans", "canonical_source": "https://forum.level1techs.com/t/quad-r9700s-on-am4-shenanigans/257456?page=4#post_71", "published_at": "2026-10-03 07:21:33+00:00", "updated_at": "2026-10-03 09:37:41.280300+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-chips", "mlops"], "entities": ["AMD Radeon R9700", "AM4", "DeepSeek", "vLLM", "llamaserver", "MXFP4", "ROCm Validation Suite", "TransferBench"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/quad-r9700-s-on-am4-shenanigans", "markdown": "https://wpnews.pro/news/quad-r9700-s-on-am4-shenanigans.md", "text": "https://wpnews.pro/news/quad-r9700-s-on-am4-shenanigans.txt", "jsonld": "https://wpnews.pro/news/quad-r9700-s-on-am4-shenanigans.jsonld"}}