2x CMP 170HX for LLM inference: unlock, cross-root P2P, serving (vLLM, EPYC/ROMED8-2T) A developer documented running two unlocked NVIDIA CMP 170HX mining cards (GA100, A100-class) for LLM inference on a single AMD EPYC 7443 / ASRock Rack ROMED8-2T host, unlocking them to 64 GB and 74 SM and getting cross-root-complex P2P working. The revised setup uses amoghmunikote's cmpunlocker patched with three P2P patches from bayley/cmpunlocker, a 300 W VBIOS cross-flash with undervolting under a power cap, and serves a W4A16 MoE model with vLLM 0.30.0 built from source. The writeup notes the initial P2P configuration was not reproducible for some users and details the UEFI settings, kernel parameters and cooling needed to make the 64 GB BARs usable. Running two NVIDIA CMP 170HX GA100, A100-class mining cards for LLM inference on a single host: unlocking them to 64 GB and 74 SM, getting P2P working across separate root complexes, cross-flashing the 250 W card to NVIDIA's 300 W VBIOS, undervolting under a power cap, and serving a W4A16 MoE model with vLLM. Update 2026-10-02. The initial setup here was not reproducible for some people P2P did not come up . This is the revised setting: amogh's cmpunlocker as the base, patched selectively with three P2P patches from bayley/cmpunlocker, pinned to exact commits section 2 . Also new: 74 SM, the 300 W VBIOS cross-flash, undervolt, thermal limits and benchmarks. - 2x NVIDIA CMP 170HX 8 GB locked, 64 GB after unlock , same board PN 900-11001-0108-000, GPU 20C2-105-A1 . - AMD EPYC 7443, ASRock Rack ROMED8-2T. The two PCIe x16 slots are on separate root complexes nvidia-smi topo -m shows NODE , which matters for P2P section 2 . - Plenty of host RAM. The MoE model in section 6 offloads its PLE / n-gram tables to system memory about 128 GB , so this is a real requirement, not just headroom. We run 256 GB DDR4 8x 32 GB and budget roughly 180 GB for the serving container. - The cards are passive. Ours each have one Same Sky formerly CUI Devices CBM-7525B-145-494-22 blower 75x75x25 mm on a straight shroud, driven at 100% duty via the BMC, case open; without forced airflow they overheat within minutes. This is our current blower solution and it is not optimal: all power and temperature limits in this gist section 4, Benchmark are for this setup. Next we will try stronger blowers, Delta BCB0812UHN-TP09 12 V, 30 CFM or Delta BFB1012UH-BA40ZYD 14.2 V nominal, 37 CFM ; both draw over 2 A, so check what your fan headers can supply. | Component | Version | Note | |---|---|---| | NVIDIA open kernel modules | 610.57.04 | built and patched by cmpunlocker | | CUDA Toolkit | 13.3 | needed for vLLM's flashinfer JIT | | cmpunlocker | amoghmunikote/cmpunlocker master 6c442ee "Add more SMs", PR 55 | not tagged yet; last release v0.4. Plus the three P2P patches from section 2 | | vLLM | 0.30.0 | built from source | | Kernel | 7.0.14-pve Proxmox VE | pin it, or re-run the installer after a kernel upgrade | | VBIOS | 92.00.6D.00.0A 300 W on both cards | section 3 | The ROMED8-2T is UEFI-only. These make the unlocked 64 GB BARs usable: - Above 4G Decoding: Enabled. Each unlocked card exposes a 64 GB BAR two cards is about 128 GB of BAR space . The firmware has to map that in the 64-bit MMIO window above 4 GB, otherwise the large BARs cannot be assigned and the driver fails. This is the setting that matters most. - Resizable BAR: Enabled Auto . Lets each card grow its BAR to the full 64 GB after the unlock. Pre-unlock the BARs are small, so this only matters afterwards. - Secure Boot: Disabled. The cmpunlocker kernel modules are unsigned. - Boot in pure UEFI mode disable CSM if the board has it . Legacy/CSM boot can leave GPU BARs unassigned symptom: "BARx is 0M" . A pci=realloc kernel argument is a stopgap; UEFI-only is the clean fix. - IOMMU: off amd iommu=off iommu=off , no AMD-Vi in dmesg, /sys/class/iommu empty . We use the cards from LXC containers, which needs no IOMMU, and cross-root P2P works with it off. An earlier version of this gist recommended amd iommu=on iommu=pt ; that is only relevant for VM passthrough and is not what this box runs. A firmware change needs a full power-off to take effect. console=tty0 console=ttyS1,115200n8 pci=realloc quiet acpi enforce resources=lax amd iommu=off iommu=off iomem=relaxed - pci=realloc : lets the kernel reassign the large 64 GB BARs; we keep it on. - amd iommu=off iommu=off : see above. - iomem=relaxed : only needed if you use 170tune's BAR0 tools section 4 ; not needed for the unlock, P2P or serving. - console=ttyS1,115200n8 : serial console for the BMC's Serial-over-LAN. - --no-iommu and --no-passthrough are installer flags of cmpunlocker leave the kernel cmdline alone, do not bind the cards to vfio-pci , not kernel arguments. ./install.sh --profile=8gb --no-iommu --no-passthrough - Use master 6c442ee or newer: it opens RECONFIG PLM and re-enables the reserved TPCs, 70 to 74 SM per card check with torch multi processor count ; dmesg shows SM-RECONFIG GPCn . v0.4 gives 70 SM. - It builds the open-gpu-kernel-modules against your running headers 610.57.04, 615.71.09, 610.43.03, 610.43.02 supported . It removes NVIDIA DKMS modules. - Secure Boot off unsigned modules . - Cold boot full power-off after the first install and after any card or firmware change. For a driver update, e.g. v0.4 to the +4 SM commit, a plain reboot was enough here and in the field report in PR 58 . If the 64 GB or the extra SMs do not show up after a reboot, power the machine off completely. - Changing the SM count invalidates vLLM's torch.compile cache. Move ~/.cache/vllm of the service user aside before the reboot, or vLLM crash-loops. - The installer rewrites /etc/modprobe.d/cmp-pcie-gen2.conf and drops any extra regkeys, including ForceP2P from section 2. Re-add them after every install, then update-initramfs -u . - Result: each card reports 64 GB, 74 SM, and the link trains at PCIe Gen2 x16. Two cards on separate root complexes default to no P2P: cudaDeviceCanAccessPeer is 0, and the GSP capabilities report GPU NOT SUPPORTED. What works here is a regkey plus three driver patches on top of cmpunlocker. The patches are from bayley/cmpunlocker https://github.com/bayley/cmpunlocker GPL-2.0, not in amogh's repo . We use them unchanged, pinned to bayley commit 5a7bb4b : - 0011-p2p-bar1.patch https://github.com/bayley/cmpunlocker/blob/5a7bb4b7e5056306fe49e8b824787659abb19914/driver/patches/0011-p2p-bar1.patch : BAR1 P2P path bus/BIF/UVM/RM changes . - 0013-skip-mailbox-peer-preinit.patch https://github.com/bayley/cmpunlocker/blob/5a7bb4b7e5056306fe49e8b824787659abb19914/driver/patches/0013-skip-mailbox-peer-preinit.patch : stops the mailbox peer pre-registration at GPU init from blocking BAR1 P2P. - 0015-bar1p2p-readcap-override.patch https://github.com/bayley/cmpunlocker/blob/5a7bb4b7e5056306fe49e8b824787659abb19914/driver/patches/0015-bar1p2p-readcap-override.patch : restores the BAR1 P2P read capability that the bridge discovery loses. To use them with amogh's cmpunlocker we run master 6c442ee https://github.com/amoghmunikote/cmpunlocker/commit/6c442eeb6448b97c803e72b61da344a39e0a26ab : copy the three files into driver/patches/ and register them in two places. driver/build.sh , array PATCH ORDER , sets the order in which the patches are applied patch -p1 , top to bottom . Add the three at the end, after cmp-sku-mask.patch , in this order. On 6c442ee the full array then reads: PATCH ORDER= sec2-postbl-plm-ss-cfg.patch booter-verify.patch late-pma.patch bar0-pramin-clamp.patch ce-scrub-workarounds.patch persistent-sw-state.patch pcie-gen2.patch pcie-gen2-probe-retrain.patch name-string.patch bar1-resize-unlock.patch cmp-sku-mask.patch 0011-p2p-bar1.patch 0013-skip-mailbox-peer-preinit.patch 0015-bar1p2p-readcap-override.patch common/constants.yaml YAML , section unlocks: , must declare every patch that PATCH ORDER builds; tools/read-constants.py cross-checks both and the build aborts with "patch ... is built but not declared in constants.yaml" otherwise. The order of the entries there does not matter. Append: p2p bar1: patch: 0011-p2p-bar1.patch registers: {} p2p skip mailbox: patch: 0013-skip-mailbox-peer-preinit.patch registers: {} p2p readcap: patch: 0015-bar1p2p-readcap-override.patch registers: {} Then run the installer and check its log logs/install