00:00
2026-10-02
mindstudio.ai
ai-infrastructure
How to Cluster Two NVIDIA DGX Sparks for Local LLM Inference
Clustering two NVIDIA DGX Sparks over a 200 GB RoCE RDMA cable and running tensor-parallel inference with vLLM lets the pair serve models in the 150 to 160 GB range, such as DeepSeek V4 Flash, that do…