DeepSeek v4.1 Flash avg 102 tps on 4x RTX6000 pro max-q, 2.1x up from v4-flash A builder documented a dual-boot Ubuntu 24.04/Windows workstation running four NVIDIA RTX6000 Blackwell Pro Max-Q GPUs (96 GB each, 384 GB total) with an AMD Ryzen Threadripper PRO 7985WX and 512 GB of DDR5 RAM, reporting DeepSeek v4.1 Flash at an average 102 tokens per second, 2.1x faster than v4-flash. The final parts cost was $54,092 plus $816 for UPS shipping, with the build timeline running from 06/30/24 to 12/23/25. The builder noted 15A power is insufficient, recommending 20A/120V or 20A/240V (NEMA 6-20R), and warned that 0-day model support on the SM120 architecture may require assembling a custom model runtime. Recently saw Supermicro at Super Compute 2025 Overview https://www.youtube.com/watch?v=eeLkbJz0r5g and thought to share my build. I got a 4x RTX6000 Blackwell Max-Q node in place that I built out from my original rtx6000 ada workstation from 2024. I’ve shared this node with friends and collegues, which we mostly used for local inference, ML experiments and pyspark jobs, so I thought to put this out as a reference for them based on my earlier notes, and as I now got good six months since my last upgrade. - dual-boot ubuntu 24.04 LTS / windows workstation for local LLM inference — 4x RTX6000 Blackwell Pro Max-Q build - CPU: AMD Ryzen Threadripper PRO 7985WX 64-core - GPU: 4x NVIDIA RTX6000 Blackwell Pro Max-Q Workstation Edition 96 GB × 4 = 384 GB - RAM: 8x Kingston KSM56R46BD4PMI-64HAI 64GB DDR5 512GB - Platform: Asus WRX90E-SAGE SE - PSU Super Flower Leadex Titanium 1700W ATX 3.1 SF-1700F14HT - Case : Corsair Obsidian 1000D - Cooling : front: 8x iCUE LINK RX120 RGB 120mm, back: 2x iCUE LINK RX120 RGB 120mm, aio: Corsair H170i Elite XT 420mm RGB - Storage : 12TB Gen5 NVMe M.2 2x Crucial T710 4TB RAID0 mdadm, 1x Crucial T705 4TB , 80TB SATA 4x Seagate Exos X24 20TB RAID0 mdadm - Status: Complete timeline: 06/30/24 - 12/23/25 - Final parts cost with WA tax : $54092 +$816 UPS Thoughts: - Is 15A enough? Not really, ideally use 20A/120V or better 20A/240V NEMA 6-20R if possible. I ran from regular 15A, but GPU workloads cap the CPU clock with sudo cpupower frequency-set -u 2.20GHz - 0-day model support : could be limited. Labs would build against datacenter hopper/blackwell SM90 / SM100 and may not release the model runtime that would run on SM120. You may need to assemble your own model runtime if you want an early support for your quants on SM120. - Cost variance : RAM and Storage prices would inflate the costs if built in mid 2026. - Build difficulty : mostly straightforward, nothing unconventional. - Airflow : Obsidian 1000D gives massive clearance but there is decent distance between front fans array and stacked GPUs, I’d go for a case geometry that is tigther. I stumbled on the You can now train a 70b language model at home https://www.answer.ai/posts/2024-03-06-fsdp-qlora.html post Mar 2024 , which got me thinking about building my own ML workstation. I wasn’t doing any prior training work so I couldn’t judge the practically of all ideas there. I hadn’t put together a custom PC since I was a kid last one had an NVIDIA GeForce 7950 GX2 , so going off-the-shelf felt like the safer bet. I looked at Lambda Labs, and their salesman was kind enough to put together a reference spec: Dual 6000 Ada, 32-core CPU, 256GB RAM at $23772 . I didn’t have a way to justify that kind of outlay, and rough math put the quote at ~18% margin over May 2024 retail parts. So the goal became a custom build, spread across 2 years of bonuses. Initially: 1. Match the reference spec, but go 512GB RAM to stay at 2x VRAM if I ever expand to 4x RTX6000 Ada. 2. Newer cases were out, but I picked the Corsair Obsidian 1000D — liked its more brittle feel, and the massive clearance left for 4 GPUs. 3. AIO and cooling from Corsair for no real reason, other than guessing a popular ecosystem might mean better support for Linux. 4. Start with one RTX6000 Ada, add one per bonus till 4x. 5. ASRock WRX90 WS EVO as the base — figured a later release than the WRX90E-SAGE meant better stability so I thought 07/19/2024: all components arrived, initial assembly with single RAM DIMM will never POST. Flushing fresh bios through remote management interface would always get stuck. ASRock customer support suggested replacement. 07/25/2024: ASRock WRX90 WS EVO replacement arrived, same problem. Refunded through newegg and ordered ASUS Pro WS WRX90E-SAGE SE instead. 08/04/2024: ASUS Pro WS WRX90E-SAGE booted on the first try and took a good ~4-5 minutes. Initial build complete. full load lands between 800~900W from the wall. 10/20/2024: added second RTX6000 Ada. now can hit 1200W 08/06/2025: gpt-oss-120b just landed and would run well on 96GB VRAM with llama.cpp but trying it inside claude code behind litellm for anything practical would go quite poorly. 2bit quantization of Qwen3-Coder-480B-A35B-Instruct-GGUF https://huggingface.co/unsloth/Qwen3-Coder-480B-A35B-Instruct-GGUF works seemingly better with regards to its tool calling but with offloading, a couple of turns would take 10-15 minutes to complete. GGUF Hardware requirement says 180GB and that means 4 ada cards to get it running. Yay, okay, so second-hands ada card can be bought around $5k each, bringing this to rougly $10k investment. At the same time, new blackwell GPUs were recently released but are completely sold out across all distributors. The catch is that the VRAM doubles per card, so 96GBs. It’s MSRP shows $8.4k and technically if I can sell current GPUs at a reasonable price, and blackwell stock stabilizes, maybe I can double the GPU VRAM and upgrade the generation for a smaller outlay then 10k. 08/30/2025: RTX6000 pro max-q blackwell https://www.newegg.com/nvidia-900-5g153-2200-000-rtx-pro-6000-96gb-graphics-card/p/N82E16814132105 began listing again at $9k, I was impatient and got it as soon as I saw, but just ~10 days later PNY began listing at MSRP of $8.4k. With new GPU installed I needed to sell old gen before the prices drop, the deal I got on itsworthmore https://www.itsworthmore.com/ was 4k per card 70% from MSRP and so once complete, I bought the second PNY RTX6000 pro max-q blackwell https://www.newegg.com/p/N82E16888884003 at MSRP. Result: doubling VRAM with total outlay of $9.4k excluding tax. 11/15/2025: various model checkpoints sitting around would use most of the storage space - ordered 4x Seagate Exos X24 20TB https://www.amazon.com/dp/B0CN5LH117?tag=level1techs-20 for the cold storage. 11/26/2025: ordered remaining two cards as we needed more compute for the coursework. Also thought to upgrade the original 1600W PSU to 1700W as 4 GPUs system would be hitting very close to that ceiling. Doing some exploration around long-running agents and running MCTS-style moatless-tree-search https://github.com/aorwall/moatless-tree-search burned a lot of tokens and would still take days to iterate agianst benchmarks. 12/22/2025: GPUs and PSU arrived and 4 GPU is complete. I no longer have a space for anthenas bracket, so I got mobile antennas WiFi 6E Antenna U.FL MHF4 Tri-Band https://www.amazon.com/dp/B0CJDVZYWF?tag=level1techs-20 that would just stick out the back case opening. final: 1. ASUS Pro WS WRX90E-SAGE SE: $1,433 amazon 2. NVIDIA RTX6000 Blackwell Pro Max-Q Workstation Edition 4x: $9954 Aug 2025 NVIDIA + $9172 Sept 2025 PNY + 2x $8853 Dec 2025 PNY newegg 3. AMD Ryzen Threadripper PRO 7985WX 3.2 GHz 64-Core sTR5 Processor: $8129 b&h photo-video 4. Kingston 64GB ECC Registered DDR5 5600 PC5 44800 Server Memory Model KSM56R46BD4PMI-64HAI 8x: $2338 at 6/30/2024 newegg 5. Crucial T705 4TB PCIe Gen5 NVMe M.2 SSD: $595 amazon 6. Crucial T710 4TB PCIe Gen5 NVMe M.2 SSD x2: $882 amazon 7. Seagate Exos X24 20TB x4: $1896 amazon 8. Super Flower Leadex Titanium 1700W ATX 3.1: $639 newegg 9. CORSAIR XTM70 Extreme Performance Thermal Paste, 3g: $43 newegg 10. Corsair Obsedian Series 1000D: $579 amazon 11. Corsair H170i Elite XT 420mm RGB Liquid CPU Cooler: $300 corsair 12. Corsair iCUE LINK RX120 RGB 120mm 10x: $334 corsair 13. Silkland 80Gbps DisplayPort Cable 2.1 6.6FT/2M: $22 amazon 14. Cable Matters 2-Pack 6 Pin PCIe to Molex Power Cable 6 Inches: $9 amazon 15. SSSUWP Motherboard 9pin USB2.0 to Dual 9pin Extension Cable x2: $14 amazon 16. SZSDMY USB 3.0 Motherboard Header Splitter,19/20 Pin 1 to 2: $15 amazon 17. Intel AX210NGW for Laptop NIC Wi-Fi 6E Wireless NIC Supports 6GHz Band Also Supports Bluetooth 5.3: $22 amazon 18. WiFi 6E Antenna U.FL MHF4 Tri-Band 6GHz 5GHz 2.4GHz Internal Laptop Wi-Fi Antennas for AX200 AX210 1675X M.2 NGFF Card $10 amazon 19. GOLDENMATE 2000VA/1600W 460Wh LiFePO4 , AVR, Line Interactive Sinewave UPS: $816 amazon Final system runs dual boot Ubuntu 24.04 LTS and Windows 11. Windows is corp enrolled, BitLocker is required and so it takes its own Crucial T705. Other drives exclusive to linux: 2x Crucial T710 4TB in RAID0 and 4x Seagate Exos X24 20TB in RAID0 - both through mdadm. Usage: Ubuntu: predominantly as a headless compute node for LLM inference sudo systemctl isolate multi-user.target that for now runs DeepSeek-V4-Flash per DEPLOY-MXFP4-W4A4-DEEPSEEK-V4-FLASH-SM120.md https://github.com/ambientlight/rtx-pro-6000-bench/blob/main/docs/DEPLOY-MXFP4-W4A4-DEEPSEEK-V4-FLASH-SM120.md . System runs stably with uptimes generally in 35 - 40 days range. Windows/WSL2: sporadically used for work as building large monorepos on 7985WX would be several times faster then on my M2 Max macbook. I didn’t get a fully headless workflow there as I needed interactive auth often, generally wasn’t able to get long multi-week uptimes as Windows Remote Desktop on my build would seldomly crash windows into green screen win11 insiders during connection attempt after several days of uptime. With 4GPUs, upgraded Superflower Leadex Titanium 1600W https://www.amazon.com/dp/B00SKAV0UQ?tag=level1techs-20 to Superflower Leadex Titanium 1700W ATX 3.1 https://www.super-flower.com.tw/products-detail/100/ . With 2 12V-2x6 sockets, there is enough to drive all 4 GPUs, but there is no spare PCIe 8pin cable to drive 2 iCUE Links back and front fans are through its own iCUE link from last available 9th 8-pin socket. This Leadex Titanium gen rolls 9 same modular 8-pin outputs but comes with 2x CPU and 6x PCIe 6+2 cables: 6x PCIe are all taken by 2 remaining GPUs + 2 auxilary PCIE 8P 1 PWR / PCIE 8P 2 PWR power that WRX90 needs for multi-GPU configuration. The keying here is different from prev gen so its Gen 5 12VHPWR PSU Cable https://www.newegg.com/p/1W7-00TK-00001?Item=9SIAMNPJSX6894 won’t work . There is one 6pin to 4x Molex in the box, but keying for SATA / Molex are the same with old gen, so I was able to reuse one molex cable from the old PSU to drive each ICUE link through seperate cable via 2 Molex → 6 Pin PCIe https://www.amazon.com/Cable-Matters-2-Pack-Molex-Inches/dp/B01DV1Z22K?tag=level1techs-20 . One other 6pin used to power 4x SATA and one used to power Commander Core that comes with AIO. This leave the PSU with one spare 6pin and one 8pin output per drawing below. WRX90 comes with 2 USB2.0 headers, each uses a splitter to wire 2 ICUe Links, AIO CPU LCD and Commander Core. I was never able to figure out how to chain built-in 1000D Commander Pro that comes with the case through this so that all 5 get detected. I only lose case’ front pannel leds so I never resolved that. Base power estimate on full load sums to rough 1760W DC, which sits above TDP, cybernetics measurements https://www.cybenetics.com/evaluations/psus/2902/ point to 90% efficiency at 100% load on 115V which would trip 15A breaker ~1950W AC wall if that load is sustained. I got to trigger it once when stress testing uncapped load when CPU boosted to 350W during full GPU load. My rental space doesn’t expose me a 20A breaker, so this runs from regular 15A behind a UPS. I needed to compromise a bit, where at default I would limit CPU Threadripper PRO 7985WX frequency upper bound in linux with sudo cpupower frequency-set -u 2.20GHz that would have CPU peaking at 190-200W with the build running stably below 1800W AC on full GPU load. All peripherals except the node itself were routed through a seperate kitchen breaker, where the node runs through the GOLDENMATE 2000VA/1600GW line-interactive sinewave UPS . I’ve got their Online Double Conversion https://www.amazon.com/dp/B0FR4381BK?tag=level1techs-20 model first but that it would trip my breaker immediately as soon as I plug into my PSU when switched off , I switched to line-interactive sinewave https://www.amazon.com/dp/B0DZ2GS9MT?tag=level1techs-20 , I cannot confirm yet if the transfer time is truly 4-6 ms per model spec, the system was stable during 2 power outages during last 6 month but it was not at a full load. For now though, UPS manages to operates over it TDP generally registering in range of 1650~1760W AC on UPS on sustained full GPU loads disregarding its over-TDP beeping . If I could I would size it against Super Flower Leadex Titanium 2200W https://www.newegg.com/super-flower-leadex-platinum-atx3-1-atx3-0-compatible-2200w-cybenetics-platinum-power-supplies/p/1HU-024C-000A6?msockid=01843f14a6876643316b2986a7ac67e1 if I had a NEMA 6-20R outlet. jurkovic-nikola/OpenLinkHub https://github.com/jurkovic-nikola/OpenLinkHub works out of the box in Ubuntu. Ran 10min gpu-burn at both 250W and 300W caps, GPU fans on auto, front case fans on sustained 2000RPM back / AIO on normal at ambient 21°. gpu burn covers tf32, so for the precisions that are used in inference bf16/fp8/fp4 I used scaled mm for fp8, flashinfer’s mm fp4 cutlass backend, the Sm120B12x kernel for fp4. GPU layout: top to bottom GPU3 bus: E1:00 position: 3 : PNY 300W cap GPU0 bus: 01:00 position: 2 : NVIDIA 300W cap GPU2 bus: C1:00 position: 1 : PNY 325W cap GPU1 bus: 02:00 position: 0 : PNY 325W cap GPUs are housed with slight gaps in between which made some difference. My previous setup had GPU3 and GPU0 swapped with sandwitched top card position 2 unable to maintain 300W after 3 minutes while thermally throttling and dropping to 272W as it reached and maintained 92°. Swapping the top 2 cards and housing them with wider gap helped to maintain 300W draw across all 4 cards during 10min burn. Note, bottom 2 PNY cards that were bought in Nov 2025 show 325W max power cap with their serial starting with 1792725 vs 1792425 for earlier 300W cap cards . Below: dense throughput broken out per card, shown as TFLOPS @ sustained-MHz / peak°C so you can see each die’s clock and temp, plus MFU vs NVIDIA’s published Max-Q dense peak. 250W | precision | spec peak/card | GPU0 | GPU1 | GPU2 | GPU3 | total | MFU | | tf32 | 219.5 | 93 @1108/89° | 101 @1177/82° | 102 @1192/89° | 107 @1194/86° | 404 | 46% | | bf16 | 438.9 | 197 @1136/89° | 216 @1203/83° | 217 @1215/89° | 227 @1222/87° | 857 | 49% | | fp8 | 877.9 | 392 @1149/89° | 426 @1211/83° | 432 @1231/89° | 451 @1240/87° | 1700 | 48% | | fp4 | 1755.7 | 697 @1174/89° | 750 @1266/83° | 766 @1260/89° | 800 @1267/87° | 3013 | 43% | 300W | precision | spec peak/card | GPU0 | GPU1 | GPU2 | GPU3 | total | MFU | | tf32 | 219.5 | 117 @1331/91° | 123 @1381/88° | 124 @1392/91° | 126 @1412/88° | 491 | 56% | | bf16 | 438.9 | 245 @1324/91° | 256 @1373/88° | 258 @1382/91° | 261 @1397/88° | 1019 | 58% | | fp8 | 877.9 | 482 @1329/91° | 501 @1376/88° | 504 @1387 https://forum.level1techs.com/u/1387 /91° | 511 @1399/88° | 1998 | 57% | | fp4 | 1755.7 | 849 @1358/91° | 880 @1410/88° | 886 @1415/91° | 900 @1430/88° | 3515 | 50% |