{"slug": "2x-cmp-170hx-for-llm-inference-unlock-cross-root-p2p-serving-vllm-epyc-romed8-2t", "title": "2x CMP 170HX for LLM inference: unlock, cross-root P2P, serving (vLLM, EPYC/ROMED8-2T)", "summary": "A developer documented running two unlocked NVIDIA CMP 170HX mining cards (GA100, A100-class) for LLM inference on a single AMD EPYC 7443 / ASRock Rack ROMED8-2T host, unlocking them to 64 GB and 74 SM and getting cross-root-complex P2P working. The revised setup uses amoghmunikote's cmpunlocker patched with three P2P patches from bayley/cmpunlocker, a 300 W VBIOS cross-flash with undervolting under a power cap, and serves a W4A16 MoE model with vLLM 0.30.0 built from source. The writeup notes the initial P2P configuration was not reproducible for some users and details the UEFI settings, kernel parameters and cooling needed to make the 64 GB BARs usable.", "body_md": "Running two NVIDIA CMP 170HX (GA100, A100-class mining cards) for LLM inference on a single host: unlocking them to 64 GB and 74 SM, getting P2P working across separate root complexes, cross-flashing the 250 W card to NVIDIA's 300 W VBIOS, undervolting under a power cap, and serving a W4A16 MoE model with vLLM.\n\n**Update 2026-10-02.** The initial setup here was not reproducible for some people (P2P\ndid not come up). This is the revised setting: amogh's cmpunlocker as the base, patched\nselectively with three P2P patches from bayley/cmpunlocker, pinned to exact commits\n(section 2). Also new: 74 SM, the 300 W VBIOS cross-flash, undervolt, thermal limits\nand benchmarks.\n\n- 2x NVIDIA CMP 170HX (8 GB locked, 64 GB after unlock), same board (PN 900-11001-0108-000, GPU 20C2-105-A1).\n- AMD EPYC 7443, ASRock Rack ROMED8-2T. The two PCIe x16 slots are on separate root\ncomplexes (`nvidia-smi topo -m` shows NODE), which matters for P2P (section 2).\n- Plenty of host RAM. The MoE model in section 6 offloads its PLE / n-gram tables to system memory (about 128 GB), so this is a real requirement, not just headroom. We run 256 GB DDR4 (8x 32 GB) and budget roughly 180 GB for the serving container.\n- The cards are passive. Ours each have one Same Sky (formerly CUI Devices) CBM-7525B-145-494-22 blower (75x75x25 mm) on a straight shroud, driven at 100% duty via the BMC, case open; without forced airflow they overheat within minutes. This is our current blower solution and it is not optimal: all power and temperature limits in this gist (section 4, Benchmark) are for this setup. Next we will try stronger blowers, Delta BCB0812UHN-TP09 (12 V, 30 CFM) or Delta BFB1012UH-BA40ZYD (14.2 V nominal, 37 CFM); both draw over 2 A, so check what your fan headers can supply.\n\n| Component | Version | Note | \n|---|---|---|\n| NVIDIA open kernel modules | 610.57.04 | built and patched by cmpunlocker | \n| CUDA Toolkit | 13.3 | needed for vLLM's flashinfer JIT | \n| cmpunlocker | amoghmunikote/cmpunlocker master `6c442ee` (\"Add more SMs\", PR #55) | not tagged yet; last release v0.4. Plus the three P2P patches from section 2 | \n| vLLM | 0.30.0 | built from source | \n| Kernel | 7.0.14-pve (Proxmox VE) | pin it, or re-run the installer after a kernel upgrade | \n| VBIOS | 92.00.6D.00.0A (300 W) on both cards | section 3 | \n\nThe ROMED8-2T is UEFI-only. These make the unlocked 64 GB BARs usable:\n\n- Above 4G Decoding: Enabled. Each unlocked card exposes a 64 GB BAR (two cards is about 128 GB of BAR space). The firmware has to map that in the 64-bit MMIO window above 4 GB, otherwise the large BARs cannot be assigned and the driver fails. This is the setting that matters most.\n- Resizable BAR: Enabled (Auto). Lets each card grow its BAR to the full 64 GB after the unlock. Pre-unlock the BARs are small, so this only matters afterwards.\n- Secure Boot: Disabled. The cmpunlocker kernel modules are unsigned.\n- Boot in pure UEFI mode (disable CSM if the board has it). Legacy/CSM boot can leave\nGPU BARs unassigned (symptom: \"BARx is 0M\"). A `pci=realloc` kernel argument is a\nstopgap; UEFI-only is the clean fix.\n- IOMMU: off (`amd_iommu=off iommu=off` , no AMD-Vi in dmesg,`/sys/class/iommu` empty).\nWe use the cards from LXC containers, which needs no IOMMU, and cross-root P2P works\nwith it off. (An earlier version of this gist recommended`amd_iommu=on iommu=pt` ; that\nis only relevant for VM passthrough and is not what this box runs.)\n\nA firmware change needs a full power-off to take effect.\n\n```\nconsole=tty0 console=ttyS1,115200n8 pci=realloc quiet acpi_enforce_resources=lax amd_iommu=off iommu=off iomem=relaxed\n```\n\n- `pci=realloc` : lets the kernel reassign the large 64 GB BARs; we keep it on.\n- `amd_iommu=off iommu=off` : see above.\n- `iomem=relaxed` : only needed if you use 170tune's BAR0 tools (section 4); not needed\nfor the unlock, P2P or serving.\n- `console=ttyS1,115200n8` : serial console for the BMC's Serial-over-LAN.\n- `--no-iommu` and`--no-passthrough` are**installer flags** of cmpunlocker (leave the\nkernel cmdline alone, do not bind the cards to vfio-pci), not kernel arguments.\n\n```\n./install.sh --profile=8gb --no-iommu --no-passthrough\n```\n\n- Use master `6c442ee` or newer: it opens RECONFIG_PLM and re-enables the reserved TPCs,\n70 to 74 SM per card (check with torch`multi_processor_count` ; dmesg shows`SM-RECONFIG GPCn` ). v0.4 gives 70 SM.\n- It builds the open-gpu-kernel-modules against your running headers (610.57.04, 615.71.09, 610.43.03, 610.43.02 supported). It removes NVIDIA DKMS modules.\n- Secure Boot off (unsigned modules).\n- Cold boot (full power-off) after the first install and after any card or firmware change. For a driver update, e.g. v0.4 to the +4 SM commit, a plain reboot was enough here (and in the field report in PR #58). If the 64 GB or the extra SMs do not show up after a reboot, power the machine off completely.\n- Changing the SM count invalidates vLLM's torch.compile cache. Move\n`~/.cache/vllm` of the service user aside before the reboot, or vLLM crash-loops.\n- The installer rewrites `/etc/modprobe.d/cmp-pcie-gen2.conf` and drops any extra regkeys,\nincluding ForceP2P from section 2. Re-add them after every install, then`update-initramfs -u` .\n- Result: each card reports 64 GB, 74 SM, and the link trains at PCIe Gen2 x16.\n\nTwo cards on separate root complexes default to no P2P: `cudaDeviceCanAccessPeer` is 0,\nand the GSP capabilities report GPU_NOT_SUPPORTED. What works here is a regkey **plus**\nthree driver patches on top of cmpunlocker.\n\nThe patches are from [bayley/cmpunlocker](https://github.com/bayley/cmpunlocker) (GPL-2.0,\nnot in amogh's repo). We use them unchanged, pinned to bayley commit `5a7bb4b`:\n\n- [`0011-p2p-bar1.patch`](https://github.com/bayley/cmpunlocker/blob/5a7bb4b7e5056306fe49e8b824787659abb19914/driver/patches/0011-p2p-bar1.patch) :\nBAR1 P2P path (bus/BIF/UVM/RM changes).\n- [`0013-skip-mailbox-peer-preinit.patch`](https://github.com/bayley/cmpunlocker/blob/5a7bb4b7e5056306fe49e8b824787659abb19914/driver/patches/0013-skip-mailbox-peer-preinit.patch) :\nstops the mailbox peer pre-registration at GPU init from blocking BAR1 P2P.\n- [`0015-bar1p2p-readcap-override.patch`](https://github.com/bayley/cmpunlocker/blob/5a7bb4b7e5056306fe49e8b824787659abb19914/driver/patches/0015-bar1p2p-readcap-override.patch) :\nrestores the BAR1 P2P read capability that the bridge discovery loses.\n\nTo use them with amogh's cmpunlocker (we run master\n[`6c442ee`](https://github.com/amoghmunikote/cmpunlocker/commit/6c442eeb6448b97c803e72b61da344a39e0a26ab)):\ncopy the three files into `driver/patches/` and register them in two places.\n\n`driver/build.sh`, array `PATCH_ORDER`, sets the order in which the patches are applied\n(`patch -p1`, top to bottom). Add the three at the end, after `cmp-sku-mask.patch`, in\nthis order. On `6c442ee` the full array then reads:\n\n```\nPATCH_ORDER=(\n    sec2-postbl-plm-ss-cfg.patch\n    booter-verify.patch\n    late-pma.patch\n    bar0-pramin-clamp.patch\n    ce-scrub-workarounds.patch\n    persistent-sw-state.patch\n    pcie-gen2.patch\n    pcie-gen2-probe-retrain.patch\n    name-string.patch\n    bar1-resize-unlock.patch\n    cmp-sku-mask.patch\n    0011-p2p-bar1.patch\n    0013-skip-mailbox-peer-preinit.patch\n    0015-bar1p2p-readcap-override.patch\n)\n```\n\n`common/constants.yaml` (YAML), section `unlocks:`, must declare every patch that\n`PATCH_ORDER` builds; `tools/read-constants.py` cross-checks both and the build aborts\nwith \"patch ... is built but not declared in constants.yaml\" otherwise. The order of the\nentries there does not matter. Append:\n\n```\n  p2p_bar1:\n    patch: 0011-p2p-bar1.patch\n    registers: {}\n\n  p2p_skip_mailbox:\n    patch: 0013-skip-mailbox-peer-preinit.patch\n    registers: {}\n\n  p2p_readcap:\n    patch: 0015-bar1p2p-readcap-override.patch\n    registers: {}\n```\n\nThen run the installer and check its log (`logs/install_<timestamp>.log`). The bayley\npatches were made against an older tree, so on `6c442ee` some hunks land with an offset\nor fuzz. That is expected and fine, for example:\n\n```\n[INFO]    0011-p2p-bar1.patch\nHunk #1 succeeded at 1839 (offset 14 lines).\n[INFO]    0015-bar1p2p-readcap-override.patch\nHunk #1 succeeded at 669 with fuzz 2 (offset -13 lines).\n[ OK ]  All patches applied\n```\n\nWhat must not appear is `FAILED`, `rejects` or a `.rej` file; then the patch did not\napply. The three do not touch the files of the SM patch.\nAfter the install, re-add the regkey below (the installer overwrites that file).\n\n```\n# /etc/modprobe.d/cmp-pcie-gen2.conf\noptions nvidia NVreg_RegistryDwords=\"RmForceEnableGen2=1;RMPcieLinkSpeed=0x1;RMForceStaticBar1=1;RMPcieP2PType=1;RMForceP2PType=1;ForceP2P=0x11\"\n```\n\nThen `update-initramfs -u` and reboot.\n\n- ForceP2P=0x11 is READ bit0 plus WRITE bit4. It is read at module load, which avoids an init-order problem where the per-PCI-device-ID default runs before the device ID is populated, so the CMP branch is never taken and the override stays unset (the driver then falls back to GSP caps and reports GNS). The three RM*P2P/StaticBar1 keys are historical and probably no-ops; they do no harm.\n- Not tested: whether the regkey alone, without the three patches, is enough. An earlier version of this gist said so; that was not verified.\n- Verify: `cudaDeviceCanAccessPeer` is 1 in both directions;`cuMemcpyPeer` works both\nways; with`NCCL_DEBUG=INFO` you see`Channel NN/0 : 0[0] -> 1[1] via P2P/CUMEM` in\nboth directions.\n- Result: vLLM TP=2 prefill is about 1.8 to 1.9 times faster than without P2P, on Gen2.\n\nOur two cards came with different VBIOS: `92.00.67.00.01` (250 W, memory 1458 MHz, SM max\n1410 MHz) and `92.00.6D.00.0A` (NVIDIA's 300 W \"OC mining\" image, memory 1728 MHz, SM max\n1695 MHz). Same board, same IDs (`10DE:20C2 / 10DE:1585`). In TP2 the 67 card sets the\npace, and no power limit fixes that, because it sits at its VBIOS clock maximum.\n\nWe flashed the 67 card to 6D with nvflash 5.867 (Linux), using TechPowerUp image\n[#268495](https://www.techpowerup.com/vgabios/268495/268495) (SHA1 `efad37d5…cddfb7`; its\ncode region is byte-identical to a dump of our own 6D card). Following the recipe in\n[aiultrayang/cmp170hx-vbios-flash](https://github.com/aiultrayang/cmp170hx-vbios-flash):\n\n- **One nvflash operation per card per cold boot** (Falcon). Any second call fails with\n\"Falcon In HALT or STOP state\". So: cold boot 1 =`--save` a rollback dump of each card,\ncold boot 2 = flash, cold boot 3 = verify. No`--list` or`--check` before the real call.\n- Keep the nvidia driver from loading (`install nvidia /bin/false` and the same for\nnvidia_uvm/modeset/drm in modprobe.d) for those boots.\n- Flash: `nvflash -i <idx> --log flash.log 92.00.6D.00.0A.rom` , interactively (in tmux;\nit needs a TTY), read the prompt (current 67, new 6D, same IDs) before pressing`y` .\nNo extra flags (no`--overridesub` , no`--mergeinforom` : the InfoROM version is\nidentical, 1001.0108.01.02).\n- nvflash keeps the card's identity by default: serial, BRD/OBD/OEM objects, retired pages/row remap (RPR/RRL), erase ledger, license block. The log lists each \"Preserve InfoROM ... object\". Telemetry objects (BBX, SEN, DEM and a few others) come from the image file, i.e. from the donor card. No effect on serving.\n- A failed write means an external SPI programmer (CH341A on the SOP8 chip). Have one at hand and keep the rollback dump.\n- Result: both cards 6D, 64 GB unlock unaffected, memory 1728 MHz, max 300 W.\n\n- Cap each card to what your cooling can hold: `nvidia-smi -pl <W>` , address cards by\nUUID rather than index, and persist it with a systemd oneshot. A module reload resets\nthe limits to the VBIOS default (250 W), so re-apply after any rmmod/modprobe. With 6D\nthe hardware maximum is 300 W.\n- Our thermal test (both cards under gpu-burn at once, blowers at 100%, open case): 200 W plateaus at 77 C, 225 W at 83-84 C (no margin), 250 W runs away (+6 C/min, no plateau). We run 200 W. These numbers reflect our non-optimal cooling (one blower per card), not a limit of the card itself.\n- Undervolt via NVML: the 6D VBIOS allows a GPC VF offset of -1000..+1000 MHz\n(`nvmlDeviceSetGpcClkVfOffset` ; the unit is MHz, a shift of the V/F curve, not mV).\nWe keep the power cap as the governor and add the offset, so the same watts buy more\nclock. Under a sustained bf16 GEMM at 200 W: +0 about 1385 MHz, +200 about 1500 MHz,\n+250 about 1525 MHz. Both +200 and +250 passed full-VRAM pattern sweeps and a bit-exact\nGEMM check with no Xid.\n- +250 hung one card within seconds under real vLLM load (Xid 175 GSP RPC timeout, then\nXid 154 \"recovery action PF FLR\"; it needed a reboot). The synthetic checks did not\ncatch it; [170tune](https://github.com/cachenetics/170tune) documents the same: a passed\ngate is necessary, not sufficient. +200 is our serving point: clean under real load,\nabout 1600-1650 MHz at 200 W.\nThe offset is volatile (gone after a reboot or driver reload); re-apply it at boot.\n- 170tune caveat: its apply path sets `nvidia-smi -pl 300` and its reset path`-pl 250` , hardcoded. On a card that cannot dissipate that, use only its test tools or\npatch those values.\n- 170tune's HBM clock and timing levers need extra FBPA PLMs opened (FBPA_MEM, FBPA PLL).\namogh's cmpunlocker does not open them (a PR for that was declined), so `170tune preflight` reports the masks as closed. The SM undervolt does not need them.\n- In LXC: the `/dev/nvidiaN` minor numbers are not stable across boots and do not match\nthe nvidia-smi index. Bind all of them into the container and select the cards with`CUDA_VISIBLE_DEVICES=GPU-<uuid>` . Symptom of the wrong card: an OOM saying\n\"total capacity of 11.63 GiB\" (that was a 12 GB card, not a lost unlock).\n\nIn-band `ipmitool raw 0x3a 0xda` (get) / `0xd6` (set duty) / `0xd8` (manual mode), the\nAST2500 command set, duty 0-100 on the ROMED8-2T. The `0x3a 0x05/0x06` from the ASRock\nFAQ returned \"invalid\" on these boards. The manual setting survives reboots and BMC power\ncycles.\n\n```\n--tensor-parallel-size 2 --enable-expert-parallel\n--gpu-memory-utilization 0.90 --max-num-seqs 16\n--disable-custom-all-reduce\n--compilation-config '{\"mode\":0,\"cudagraph_mode\":\"FULL_AND_PIECEWISE\"}'\n--speculative-config '{\"method\":\"mtp\",\"num_speculative_tokens\":4}'\n```\n\nwith `VLLM_PLE_CPU_OFFLOAD=1` and `NCCL_P2P_LEVEL=SYS`.\n\n- VLLM_PLE_CPU_OFFLOAD=1 keeps the PLE / n-gram tables in host RAM (about 128 GB; budget roughly 180 GB for the container/host). Without it you get a GPU OOM.\n- `--disable-custom-all-reduce` is required with P2P. The custom IPC all-reduce backend\nOOMs the VRAM and throws \"invalid argument\" across root complexes, which crash-loops\nthe workers. NCCL/PYNCCL still uses P2P, so the speedup stays.\n- cudagraph_mode FULL_AND_PIECEWISE gave about 10% more single-stream decode on code than FULL_DECODE_ONLY.\n- Build fix for compressed-tensors PLE: vLLM 0.30.0 throws\n`NotImplementedError: Qwen4Exp PLE embedding does not support CompressedTensorsConfig` .\nIts`from_quant_config` path handles only ModelOpt/Fp8 for excluded PLE layers. Add a\nCompressedTensorsConfig branch that checks`should_ignore_layer(...)` and returns the\nunquantized PLE embedding method. Small Python change, no recompile.\n- Does not work: `--kv-cache-dtype fp8` crashes the state-attention backend on this\nmodel (only auto/bfloat16 are supported).\n\n- flashinfer's JIT compiles on first start (about 2 minutes), then it is cached.\n- CUDA tools inside a container need the matching `libcuda.so.1` /`libnvidia-ml.so.1` on the loader path; the driver userspace has to match the kernel module version.\n- A crashing vLLM worker leaks shared memory in /dev/shm (up to the PLE size) and pinned RAM. If that happens: stop the service, reset-failed, check /dev/shm and free RAM, then start once cleanly.\n- After any driver, VBIOS or clock change, check `dmesg | grep -i xid` . Xid 13/31/43 under\nload usually means an unstable clock/voltage point; 79 means the card fell off the bus.\n\n2x CMP 170HX, vLLM 0.30 TP=2, W4A16 MoE (Qwen3.8-Flash-Next), MTP 4, both cards capped at 200 W unless noted.\n\nMethod: decode counted with the server counter `vllm:generation_tokens_total` (counting\nSSE chunks undercounts with speculative decoding), real code prompts (random tokens kill\nthe MTP acceptance), 3 reps. Prefill with `vllm bench serve` random input, output 4, 3-rep\nmedian, tok/s = input length / TTFT. Use a different `--seed` for every point and rep:\nthe random dataset builds prompts as consecutive tokens from a seed-dependent offset, so\nwith one seed the 128k prompt starts with the 65k prompt and you measure prefix-cache\nhits (we got a fake 8,600 tok/s at 128k that way). Check\n`vllm:prefix_cache_hits_total` stays flat.\n\n| setup | 8k | 32k | 65k | 128k | \n|---|---|---|---|---|\n| no P2P (L1T W4A16 reference) | ~2500 | ~2500 | ~2500 | ~2500 | \n| P2P, mixed VBIOS (67 + 6D), 70 SM | 4634 | 4572 | 4543 | 4486 | \n| P2P, both 6D, 70 SM | 4682 | 4623 | 4586 | 4527 | \n| P2P, both 6D, 74 SM | 4766 | 4708 | 4669 | 4599 | \n| **+ offset +200 MHz, 200 W** | **5033** | **4989** | **4951** | **4881** | \n| + offset +200 MHz, 225 W | 5084 | 5037 | 5002 | 4930 | \n| + offset +200 MHz, 250 W | not stable |  |  |  | \n\n| setup | code c=1 | code c=4 aggregate | code c=4 per request | prose c=1 | \n|---|---|---|---|---|\n| P2P, mixed VBIOS, 70 SM | 193 | 483 | 131 | 113 | \n| P2P, both 6D, 70 SM | 208 | 508 | 140 | 122 | \n| P2P, both 6D, 74 SM | 209 | 514 | 143 | 123 | \n| **+ offset +200 MHz, 200 W** | **223** | **527** | **150** | **130** | \n| + offset +200 MHz, 225 W | 226 | 554 | 150 | 136 | \n| + offset +200 MHz, 250 W | not stable |  |  |  | \n\n| cap | SM clock under real load | power avg (prefill) | GPU max (decode / prefill) | HBM max | gpu-burn, both cards | \n|---|---|---|---|---|---|\n| 200 W | ~1600-1620 MHz | ~192-196 W | 71 / 72 C | ~77 C | plateau | \n| 225 W | ~1610-1635 MHz | ~214-219 W | 74 / 79 C | 80 C | 84-85 C after 4.5 min, aborted at 85 C (no plateau) | \n| 250 W | n/a | n/a | n/a | n/a | 83 C within 2 min, aborted; not stable with our cooling | \n\nMTP acceptance about 80% on code, 36% on prose throughout. Decode is mostly memory-bound, so the 6D memory clock helped it and the extra SMs alone did not. Prefill is compute-bound, but both cards sit at the power cap, so +4 SM gave only about +1.8%. With the +200 offset the SMs pay off: about +7% decode and +6% prefill at the same 200 W and the same temperatures. With the offset the cards already run near their practical clock ceiling at 200 W, so 225 W adds only 1-2% for 5-7 C more. 250 W is not stable here: the cards run away thermally under sustained load. We serve at 200 W, offset +200. (+250 hung a card under real load, see section 4.)", "url": "https://wpnews.pro/news/2x-cmp-170hx-for-llm-inference-unlock-cross-root-p2p-serving-vllm-epyc-romed8-2t", "canonical_source": "https://gist.github.com/ropuls/d2978531ba3a1a183e70f62f3de2eb84", "published_at": "2026-10-01 21:00:48+00:00", "updated_at": "2026-10-02 12:37:51.440627+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-chips", "mlops"], "entities": ["NVIDIA", "CMP 170HX", "AMD", "EPYC 7443", "ASRock Rack ROMED8-2T", "vLLM", "cmpunlocker", "CUDA"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/2x-cmp-170hx-for-llm-inference-unlock-cross-root-p2p-serving-vllm-epyc-romed8-2t", "markdown": "https://wpnews.pro/news/2x-cmp-170hx-for-llm-inference-unlock-cross-root-p2p-serving-vllm-epyc-romed8-2t.md", "text": "https://wpnews.pro/news/2x-cmp-170hx-for-llm-inference-unlock-cross-root-p2p-serving-vllm-epyc-romed8-2t.txt", "jsonld": "https://wpnews.pro/news/2x-cmp-170hx-for-llm-inference-unlock-cross-root-p2p-serving-vllm-epyc-romed8-2t.jsonld"}}