# How to Cluster Two NVIDIA DGX Sparks for Local LLM Inference

> Source: <https://www.mindstudio.ai/blog/how-to-cluster-dgx-spark-vllm/>
> Published: 2026-10-02 00:00:00+00:00

# How to Cluster Two NVIDIA DGX Sparks for Local LLM Inference

How to connect two NVIDIA DGX Sparks over RDMA and run tensor-parallel inference with vLLM for models too big for one unit.

## What does clustering two DGX Sparks actually mean?

Clustering two NVIDIA DGX Sparks means connecting them with a high-speed network link and running an inference engine that splits a single model across both machines so they behave like one larger GPU system. Each DGX Spark ships with 128 GB of memory. On its own, that’s not enough for some newer open models. Connected together over RDMA with a tool like vLLM running tensor parallel inference, the two units pool their memory and compute, letting you load and serve models in the 150 to 160 GB range, like DeepSeek V4 Flash, that won’t fit on a single Spark.

## TL;DR

- **Tensor parallelism** splits every layer of a model in half, with one Spark handling each half and the two nodes exchanging results over the network before the next layer can run.
- The two Sparks connect using a **200 GB RoCE cable** (RDMA over Converged Ethernet), measured in practice at around 111 GB/s rather than its rated 200, which still lets the nodes swap data fast enough for real-time inference.
- Models like **DeepSeek V4 Flash** (roughly 150 to 160 GB) and Qwen3-Next don’t fit in one Spark’s 128 GB of memory but load comfortably once split across two, leaving each node holding a bit over 100 GB including quantized weights and context overhead.
- In raw token generation speed, the dual-Spark cluster roughly tied an Apple M5 Ultra Mac Studio on DeepSeek (about 38 tokens/sec each) but lost a bit of ground on Qwen, where the Mac’s much higher memory bandwidth (1.2 TB/s versus 273 GB/s per Spark) gave it an edge.
- Where the Spark cluster pulled far ahead was **prompt processing on large contexts** : at 32,000 tokens, the Sparks finished in about 17 seconds versus roughly 50 seconds on the Mac, a gap that widened further at 128,000 tokens.
- Multi-user throughput also favored the cluster: the Mac’s total tokens-per-second peaked around four concurrent users and declined after that, while the dual-Spark setup kept climbing as more users were added.
- Quantization format matters as much as hardware. Switching inference engines (llama.cpp versus MLX on the Mac, vLLM on the Sparks) changed output speed by double-digit percentages on the same model.

## Other agents start typing. Remy starts asking.

Scoping, trade-offs, edge cases — the real work. Before a line of code.

## How do you physically connect two DGX Sparks?

The two units link through a single high-bandwidth cable rated at 200 GB, running RoCE (RDMA over Converged Ethernet). RDMA, remote direct memory access, lets each Spark write directly into the other’s memory without routing data through the CPU the way a normal network transfer would. That matters for tensor-parallel inference because every layer of the model requires both nodes to exchange partial results before computation can continue to the next layer. Any added latency in that exchange shows up directly in token generation speed.

In practice, the cable didn’t hit its full rated throughput. Measured bandwidth came in around 111 GB/s rather than the advertised 200. That’s still enough overhead to keep the two nodes synchronized without becoming the obvious bottleneck in most of the benchmarks run, though the exact performance cost of that gap wasn’t isolated in testing.

## How does vLLM split a model across two Sparks?

vLLM is the inference engine most commonly paired with NVIDIA hardware for this kind of setup (llama.cpp is another option that runs on Sparks too, but vLLM is positioned as NVIDIA’s preferred path for splitting workloads across multiple devices). The specific technique used here is tensor parallelism: instead of giving each Spark a different chunk of the model’s layers, every layer is cut in half, with each half living on a different Spark.

For every token generated, both machines compute their half of each layer, then swap results over the RDMA link before the next layer can start. That means thousands of small, fast round trips per response, which is exactly the kind of workload RDMA is built to minimize latency for. The practical payoff: a model like DeepSeek V4 Flash, quantized down to roughly four bits so it fits with room to spare, ends up with each Spark holding a bit over 100 GB out of its 128 GB capacity, leaving headroom for context and system overhead.

## Why would you need two Sparks instead of one?

A single DGX Spark has 128 GB of memory. Several newer open-weight models exceed that when quantized, even at four-bit precision. DeepSeek V4 Flash lands around 150 to 160 GB, and Qwen3-Next’s larger variant is in a similar category. Trying to load either onto one Spark simply fails, there isn’t enough room once you account for the model weights plus the memory needed to hold context during inference.

Splitting the model across two Sparks resolves that ceiling. Each node ends up responsible for half the model’s weights and half the computation per layer. This is also why tensor parallelism specifically (rather than other parallelism strategies) made sense for this setup: it evenly distributes both memory load and compute load across identical hardware, rather than assigning whole layers or whole requests to one machine at a time.

## How does the dual-Spark cluster compare to a high-end Mac Studio?

## Remy doesn't build the plumbing. It inherits it.

Other agents wire up auth, databases, models, and integrations from scratch every time you ask them to build something.

Remy ships with all of it from MindStudio — so every cycle goes into the app you actually want.

Comparing the dual-Spark cluster against an M5 Ultra Mac Studio with 256 GB of unified memory (matching the Sparks’ combined 256 GB) revealed two distinct performance profiles rather than a flat win or loss.

On pure token generation speed with a short prompt, the two systems were close. DeepSeek ran at about 38 tokens per second on both. Qwen favored the Mac, 45 tokens per second versus 38 on the Sparks, largely attributable to the Mac’s far higher memory bandwidth (1.2 TB/s rated, versus 273 GB/s per Spark). Since token generation is memory-bandwidth bound, that gap is expected on models and configurations where the Mac’s bandwidth advantage dominates.

Prompt processing told the opposite story. That stage is compute bound, handled by raw matrix multiplication on the GPU, and the Sparks’ GPUs processed large prompts much faster. At 32,000 tokens using a real 14-file Python codebase as the test input, the Sparks finished prompt processing in 17 seconds while the Mac took about 50 seconds, a roughly 3x gap. At 128,000 tokens, the Sparks completed the job in about 72 seconds; the Mac wasn’t even tested at that length because the trend suggested it would take over three minutes.

Multi-user behavior diverged too. Total throughput on the Mac peaked around four concurrent users (about 66 tokens per second combined on DeepSeek) and dropped as more users were added (46 tokens per second at eight users), because the Mac spent more time re-reading new prompts than writing responses. The Spark cluster kept scaling upward through the tested range instead of degrading.

## Is clustering two DGX Sparks worth it?

It depends on the workload. If your use case involves long prompts, large codebases, or serving multiple concurrent users, the dual-Spark cluster’s compute-bound prompt processing advantage and better multi-user scaling make a real difference, sometimes multiple times faster than a single high-memory Mac. If your workload is mostly short prompts with long generated outputs on a model like Qwen, raw memory bandwidth matters more, and a Mac Studio with very high bandwidth can match or beat the cluster on output speed.

It’s also worth noting that in real usage with coding agents or chat sessions, much of the prompt-processing cost only happens once per session if the server caches previous context. A 16,000-token cached context dropped the Mac’s follow-up response time to about 3 seconds and the Sparks’ to about 1.4 seconds, both far faster than a fresh cold start. That narrows the practical gap for ongoing conversations, even if the first message in a new session still favors the Sparks heavily.

Cost is a factor too. Two Sparks at their current pricing (climbing toward roughly $5,000 each) plus a RoCE cable land in a similar range to a comparably configured Mac Studio, so the decision comes down to workload shape rather than a clear price advantage either way.

## Frequently Asked Questions

### What is RoCE and why does it matter for clustering GPUs?

## One coffee. One working app.

You bring the idea. Remy manages the project.

RoCE stands for RDMA over Converged Ethernet. It lets two machines write directly into each other’s memory over a network connection without routing through the CPU, which reduces latency. That’s important for tensor-parallel inference, where the two Sparks must exchange partial results after every layer of computation for every token generated.

### Can you run models on a single DGX Spark without clustering?

Yes, for models that fit within 128 GB after quantization. Clustering becomes necessary once a model’s size, plus the memory needed for context, exceeds what one Spark can hold, which is the case for larger models like DeepSeek V4 Flash and Qwen3-Next’s bigger variants.

### Does vLLM work on both DGX Sparks and other hardware?

vLLM is a general-purpose, memory-efficient inference and serving engine for large language models and is commonly used on NVIDIA hardware, including DGX Sparks, specifically because it supports tensor parallelism for splitting models across multiple devices. Other tools like llama.cpp can also run on Sparks, but vLLM is the setup NVIDIA’s own hardware is typically configured to use for this kind of multi-device split.

### Why is prompt processing faster on Sparks but token generation sometimes faster on a Mac?

Prompt processing (also called prefill) is compute bound, dominated by matrix multiplication on the GPU, which favors the Sparks’ GPU horsepower. Token generation (decode) is memory-bandwidth bound, since each token requires reading the model’s full set of weights from memory, which favors hardware with higher memory bandwidth, like the Mac Studio’s 1.2 TB/s rating compared to each Spark’s 273 GB/s.

### Does the inference engine choice change performance on the same hardware?

Yes, significantly. On the same Mac Studio, switching the DeepSeek model from llama.cpp to MLX increased output speed by roughly 34% in testing. Engine choice can matter as much as the underlying hardware for certain workloads.
