Nvidia’s ability to sell GPUs is ultimately limited by how much juice the power grid can provide. With ever-growing depreciation cycles, it’ll be years before datacenters decommission their aging Hopper or Blackwell systems. More GPUs mean pulling more power from the grid.
Nvidia can’t exactly force grid operators to add capacity any faster, but it can make it easier for its customers to build smarter and more efficient bit barns.
“At the datacenter scale and at the AI-factory scale, we're literally trying to think about how can we eke out every bit of efficiency to drive more performance per gigawatt,” Dion Harris, senior director of Nvidia HPC and AI Hyperscale Infrastructure Solutions, told El Reg in a recent interview.
At the AI Infra Summit this week, we got our first look at the systems Nvidia has been building to maximize the amount of power available for compute while minimizing the impact of datacenters on the local grid.
Datacenters, as a general rule, rarely operate anywhere close to the peak capacity. A 100 megawatt datacenter might use at most 80 percent for critical compute loads. The actual ratios vary from bit-barn to bit-barn, but this provides a buffer for hardware inefficiency, conversion losses, and other spikes in demand. The downside, of course, is that this leaves 20 megawatts or so of untapped capacity that the datacenter can’t use and the utility can’t reclaim.
Every kilowatt of stranded power is a GPU that Nvidia could have sold, so Nvidia’s DSX platform aims to address both problems.
Minimizing overheads
The first of these, which Nvidia calls DSX MaxLPS, is an evolution of an old idea. If the compute and all the physical infrastructure — power cabinets, batteries, coolant distribution units (CDUs), and chillers — can talk to one another, operators can achieve significant power savings.
For example, if the air handlers had a way of knowing how much power a rack was pulling, they could ramp up and down based on demand rather than running maxed out all the time. The challenge, as you might expect, is getting systems from dozens of different vendors to speak the same language. So while the idea sounds great in theory, getting everyone on the same page was easier said than done.
That changed with the widespread deployment of AI systems, Harris explained. More efficient bit barns can churn out more tokens, which translate into higher revenues — so long as people are willing to pay for the tokens anyway.
However, Nvidia still needed to address the communications layer, something that was no doubt made easier by the fact its GPUs are the hottest commodity in the world right now.
“DSX Exchange is really kind of an API that allows us to capture information not just around the core systems,” Harris explained. “We can capture information from the other DSX ready providers that provide assets. That would be like Vertiv and Schneider Electric, and all the building management systems."
At the AI Infra Summit this week, Nvidia and Lambda showed how this bridge, working as part of its broader DSX MaxLPS offering, could be used to pack more compute into the same power envelope. In testing, Lambda was able to cram 19 nodes into the same power budget normally occupied by 16, boosting cluster-wide throughput by 24 percent and performance per watt by a similar margin.
The result, the company claims, is that datacenters don’t need to overprovision their bit barns to the same extent since they have greater visibility and control over the datacenter as a whole.
Shedding the load
Along with helping its customers get more out of their bit barns, Nvidia also showed how its datacenter management tech could be used to minimize the impact of large AI training and inference deployments on the local grid.
Working with Emerald AI and Silicon Valley Power, Nvidia demonstrated how its DSX Flex platform could be used to free up datacenter power when grid demand spikes without disrupting critical workloads.
But just like Nvidia’s DSX MaxLPS offering, the underlying tech isn’t exactly new. Demand response has been around for years now and allows utilities to ask power hungry industries to curb their energy use during periods of peak demand.
Google and others have been toying with this tech for some time now. You may recall last year when it announced it would non-essential AI workloads in order to avoid over the grid.
Nvidia’s DSX Flex aims to bring this capability to anyone deploying its hardware. But this tech may be less about keeping AI from causing brownouts and more about getting utilities to green light additional capacity on the proviso that they can reclaim some portion of it at a moment's notice.
“It allows, in some cases, the grid providers, transmission line owners, to be a little a little bit less conservative in how they allocate or over provision because now, knowing that you have the ability to curtail within a certain window, they don't have to have as much headroom,” Harris explained.
But just because a utility claws back capacity doesn’t necessarily mean that the workloads shut down, he notes. While critical workloads keep running, deep integration with Nvidia's software stack means that non-essential workloads can either be d or migrated to neighboring datacenters with excess capacity — a concept sometimes described as chasing the sun.
Another walled garden
As you probably already guessed, while Nvidia’s DSX platform works great for AI factories built using validated hardware, things get a bit more complicated as third-party accelerators from the likes of d-Matrix, SambaNova, and others are added to the mix.
For partners, like d-Matrix already living inside its NVLink Fusion ecosystem, Harris sees a path forward for extending support for DSX to these platforms, though he notes that some degree of software integration will be required and not every chip will expose the same level of granular control as its own chips do. When it comes to competing platforms, like AMD’s Instinct GPUs, bit barn builders will likely need to look to alternative datacenter management systems to replicate DSX’s capabilities.
So, on top of making the most of the limited grid capacity available today, Nvidia’s DSX is another walled garden that ensures its customers continue buying its equipment. ®