# The M5 Ultra Mac Studio might be the cheapest wall plug in local AI

> Source: <https://tokenstead.ai/guides/m5-ultra-power-efficiency-dgx-spark-rtx>
> Published: 2026-08-30 14:21:29+00:00

## The claim

A widely shared post from Jun Song (@jun_song, 44.7K views, August 29, 2026) makes the efficiency case with three real-world power numbers: a Mac Studio M5 Ultra with 512GB of unified memory draws about **350 Watts** at load, a cluster of four DGX Spark units draws about **600W**, and a six-card RTX 6000 Pro rig draws about **3900W** - more than 10x the Mac. The argument: the Mac produces faster token output than a DGX Spark while pulling roughly half of the Spark cluster’s power, and although the RTX rig beats it on raw throughput, “normal users simply don’t need that kind of throughput.” Enterprise batch inference at TokenMaxxing scale is where the RTX build earns its watt bill.

The direction of the claim matches what this site has argued since the Spark shipped: see [512GB local AI: M5 Ultra vs DGX Spark cluster vs Strix Halo](/guides/512gb-local-ai-m5-ultra-vs-dgx-spark-cluster-vs-strix-halo) and [self-host your AI: DGX Spark, RTX builds, and Mac](/guides/self-host-ai-dgx-spark-rtx-mac). Power is the dimension nobody quotes in the model-fitting arguments, and it is the one you pay for every month.

## What checks out, and what is not measured yet

The honest caveat first: **the M5 Ultra 512GB has not shipped** (units arrive September 22, with the 512GB config later), so the ~350W figure is an estimate, not a meter reading. But the surrounding numbers make it plausible. Apple’s own figures for the previous-gen M3 Ultra Mac Studio top out at 270W max system draw, and independent wall-meter testing of the M5 Ultra 256GB - same quad-die chip, half the memory - measured **34W idle and roughly 214W peak** sustained under batch inference. 350W under heavy load on the 512GB config is a credible high-side estimate, not a fantasy.

The comparison rows hold up well against published specs:

| Rig | Claimed draw | Sanity check |
|---|---|---|
| Mac Studio M5 Ultra 512GB | ~350W | M3 Ultra officially maxes at 270W; M5 Ultra 256GB independently measured ~214W peak. Plausible. |
| 4x DGX Spark | ~600W | About 150W per box under LLM load against a 170W chip - right where the
|

The bandwidth arithmetic behind “faster than a DGX Spark” also holds: the M5 Ultra’s ~1.2 TB/s of memory bandwidth beats even four Sparks pooled together (~273 GB/s each, ~1.1 TB/s aggregate), and single-GB10 throughput on proven models ranges from about 12 tok/s (dense models at FP16) to around 200 tok/s (small MoE on the NVFP4 path). The Mac’s measured aggregate - roughly 1,700+ tok/s at batch 8 on a 70B-class model - is far past any single Spark and comfortably past the cluster on real workloads.

## Tokens per watt, honestly

Where the efficiency framing deserves one correction is duty cycle. The independent M5 Ultra metering found **8.2 tokens per Watt at batch 8 sustained** - but the same testing tagged a **34W idle floor**, high for a desktop because of the quad-die design. At low utilization the Ultra’s blended efficiency drops toward half that of an M5 Max laptop pulling the same workload. The Ultra wins on efficiency when it is saturated; an interactive hobbyist poking at models a few hours a week may be better served by smaller silicon.

The tweet actually concedes this correctly. RTX 6000 Pro rigs exist for batch throughput: if your workload is “burn the whole corpus through the model,” 3900W is the price of the fastest wall plug, and amortized over millions of tokens the electricity line item changes per-cost math less than you would think. But per-device, **nothing in the unified-memory consumer class touches the Mac’s tokens-per-Watt**, and that is the comparison most self-hosters are actually making.

## What it means for buyers

If you are choosing between a Mac Studio and a memory cluster this cycle, put power on the spreadsheet next to memory bandwidth. A single M5 Ultra box replaces a wiring question (one outlet versus six cards’ worth of circuits and their heat) with a 350W line item, and it runs the [open-weight models in our catalog](/models) at the throughputs above. Run your own workload through the [rig finder](/find); the electricity math is the same one we applied to cloud-versus-local in [running GLM-5.2 locally](/guides/run-glm-5-2-locally), just in reverse - here it decides which local rig to buy.
