{"slug": "how-nvidia-dsx-maxlps-maximizes-ai-factory-throughput-and-efficiency", "title": "How NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency", "summary": "NVIDIA DSX MaxLPS, a policy-governed dynamic power-sharing system, lets AI data centers deploy up to 40% more GPUs within the same approved power budget, according to a joint NVIDIA and Nscale evaluation run on NVIDIA GB300 NVL72 systems with Kimi K2.5 workloads at Nscale's Verne campus data center in Keflavík, Iceland. The evaluation used NVIDIA Blackwell Ultra GPUs, Kimi K2.5 in FP4, NVIDIA Dynamo, NVIDIA TensorRT LLM, an 8K input sequence length and a 1K output, and measured the power-versus-performance trade-offs of reallocating stranded static power reservations across participating nodes.", "body_md": "Every unused watt is capacity left on the table. AI factories are typically provisioned for the unlikely moment when every GPU reaches peak power, creating a protective buffer that can leave valuable infrastructure underused during normal operation. NVIDIA DSX MaxLPS uses policy-governed power sharing to allocate power across participating resources dynamically, enabling customers to deploy up to 40% more GPUs within the same approved power budget.\n\nThis technical walkthrough examines a joint NVIDIA and Nscale evaluation of this approach with Kimi K2.5 workloads running on NVIDIA GB300 NVL72 systems at Nscale’s data center at the Verne campus in Keflavík, Iceland, powered entirely by renewable energy. It covers the measured trade-offs between power and performance, explains the controls used to maintain electrical limits, and presents a repeatable validation method operators can use before deploying at scale.\n\n## How static provisioning leaves usable power stranded\n\nAn AI factory operates within a hierarchy of electrical limits. Utility service, substations, power-distribution equipment, racks, nodes, and GPUs all impose constraints. Operators must keep each managed boundary within its approved limit while meeting application throughput and latency objectives.\n\nStatic power planning typically reserves enough power for every node to reach its specified peak at the same time. This conservative approach is straightforward, but AI workloads rarely draw constant power. Training workloads move through compute, communication, synchronization, and checkpointing. Inference workloads alternate among prefill, decode, memory-bound work, network activity, and idle intervals. Even instances of the same model can draw different power as request shapes and concurrencies change.\n\nThis variability creates a gap between reserved peak power and actual consumption. With static per-node reservations, unused capacity inside one reservation cannot be applied to another node. The aggregate facility can remain below its limit while additional GPU capacity stays offline.\n\n[DSX MaxLPS](https://docs.nvidia.com/dsx/maxlps/overview) monitors actual power consumption and dynamically reallocates available power across participating resources while preserving the operator’s aggregate budget and policy boundaries.\n\n## Inside the DSX MaxLPS control loop\n\nDSX MaxLPS combines chip, system, thermal, and software technologies to maximize AI factory output within land, power, and shell (LPS) constraints. Dynamic Power Software provides the control layer for policy-governed power allocation.\n\nThe control process has five technical elements:\n\n- Topology and resource groups. Operators map the participating infrastructure and organize nodes into a managed group with an aggregate power budget.\n- Telemetry. The system collects GPU, node, rack, and group power telemetry at intervals sufficient to detect available headroom and emerging power events.\n- Policy. Operator-defined rules establish node limits, group limits, allocation priorities, reserve requirements, and responses to maintenance or emergency events.\n- Allocation and control. When some resources draw less than their allocation, the software adjusts participating GPU power limits so other resources can use the available capacity.\n- Validation and enforcement. The system compares measured power against the approved group budget and adjusts allocations when consumption approaches a limit.\n\nThis is coordinated allocation, not an increase in the site’s power supply. The control loop enables more productive work beneath the same managed power budget.\n\n## How the method was evaluated\n\nFor the evaluation, Nscale deployed the MaxLPS software in its data center and collected telemetry while NVIDIA ran the workloads. The evaluation measured control behavior and workload trade-offs.\n\nThe evaluation used NVIDIA Blackwell Ultra GPUs, Kimi K2.5 in FP4, NVIDIA Dynamo, NVIDIA TensorRT LLM, an 8K input sequence length, and a 1K output sequence length. The workload mix combined high-throughput and low-latency inference instances, creating distinct power and service profiles within the managed group.\n\nThe static baseline used 35 four-GPU nodes, or 140 GPUs. It ran two high-throughput instances using 52 GPUs each and one low-latency instance using 36 GPUs. The DSX MaxLPS configuration used 48 four-GPU nodes, or 192 GPUs. It added a third 52-GPU high-throughput instance while retaining the same 36-GPU low-latency instance.\n\nJobs ran across four racks, with each distributed workload confined to a single rack in both configurations. This controlled for cross-rack performance differences. Confining distributed workloads to a single rack is not required when the test environment has been validated for equivalent performance across racks.\n\nThe team measured normalized aggregate and per-instance throughput, time to first token, end-to-end latency, interactivity, GPU, CPU, and rack power telemetry. Comparing these metrics prevents a throughput gain from hiding latency, stability, or power-compliance regressions.\n\n## What the measurements show\n\nTable 1 compares capacity, throughput, and power use for the static baseline and DSX MaxLPS configurations.\n\n| **Metric** | **Static baseline** | **DSX MaxLPS** | **Change** | \n|---|---|---|---|\n| Managed GPUs | 140 | 192 | +37.1% | \n| Aggregate throughput | 1,084,503 tokens/s | 1,618,443 tokens/s | +49.2% | \n| High-throughput output per instance | 59,153 tokens/s | 59,220 tokens/s | +0.1% | \n| Low-latency output per instance | 2,265 tokens/s | 2,265 tokens/s | 0% | \n| Mean GPU power | 97.0 kW | 131.8 kW | +35.9% | \n| Total measured power | 166.2 kW | 198.9 kW | +19.7% | \n| Power-budget utilization | 62.9% | 75.2% | +12.3 percentage points | \n| Throughput per provisioned watt | 4.10 tokens/s/W | 6.12 tokens/s/W | +49.2% | \n\n*Table 1. Static baseline and DSX MaxLPS performance and power results*\n\nBoth configurations used the same **264.4 kW provisioned power budget**. Throughput per provisioned watt divides normalized aggregate throughput by this fixed denominator. The baseline delivered 4.10 tokens/s/W, and DSX MaxLPS delivered 6.12 tokens/s/W, a 49.2% increase. Because the provisioned-power denominator remained unchanged, this percentage matches the aggregate-throughput increase.\n\nThe per-instance results remained effectively unchanged at the displayed precision. This shows that the larger managed fleet increased aggregate throughput without materially reducing the throughput of the existing high-throughput or low-latency instances.\n\nFor latency, **P75 is the 75th-percentile result**, meaning 75% of requests completed at or below that latency. **P99 is the 99th-percentile result**, representing tail behavior: 99% of requests completed at or below that latency, while the slowest 1% took longer. Median and P75 latency remained within 5% of baseline. P99 time to first token increased 17% from the 15.7-second baseline, showing the importance of evaluating tail latency alongside capacity and throughput.\n\n## Trade-offs exposed by the evaluation\n\nDynamic power allocation makes engineering trade-offs observable and controllable at the fleet level; it does not eliminate them.\n\n### Workload mix shapes available headroom\n\nDSX MaxLPS uses previously unused headroom, so added capacity still increases average power utilization. The available headroom also depends on the workload mix: complementary power profiles create more opportunity than workloads that peak simultaneously. Operators should therefore test representative production workloads against the aggregate power limit. Typical AI factories run heterogeneous workloads with different power profiles. The study therefore used a workload mix designed to reflect that variability.\n\n### Tail latency reveals service trade-offs\n\nStable median latency can mask changes in tail latency. In this evaluation, median and P75 latency remained within 5% of baseline, but P99 time to first token increased by 17%. Production acceptance criteria should be defined as part of the study.\n\n### Telemetry supports reliable control\n\nDynamic allocation depends on reliable telemetry. Missing, delayed, or incorrectly mapped measurements can undermine fleet-level decisions. For this evaluation, site-level telemetry was used to verify rack-level power measurements. Before deployment, operators should confirm that added throughput does not compromise service quality or compliance with power limits.\n\nTogether, these trade-offs define what operators should validate before deployment.\n\n## How operators can validate DSX MaxLPS\n\nUse a staged validation process with explicit boundaries and acceptance criteria.\n\n1. **Define the managed boundary.** Map utility, distribution, rack, node, and GPU topology. Set the resource-group budget, reserve requirements, and escalation behavior. Confirm which measurement represents the enforceable limit.\n2. **Establish a representative baseline.** Run the representative AI factory workload mix under static provisioning, using the intended deployment and placement rules. Measure performance and power long enough to capture workload variation and confirm repeatability.\n3. **Introduce policies conservatively.** Begin with limits close to the validated baseline. Confirm telemetry, topology, control response, and budget compliance before adding nodes.\n4. **Add capacity and test each stage.** Increase the managed population incrementally, comparing aggregate and per-instance performance at each step. Before proceeding, verify service behavior under peak demand, operating transitions, telemetry failure, and reduced power availability.\n5. **Set production operating limits.** Approve a configuration only when it meets throughput and latency objectives, stays within the managed budget, preserves the required reserve, and behaves predictably during faults and transitions.\n\nThis evaluation shows how DSX MaxLPS can reclaim stranded capacity in a power-constrained AI factory. The same policy-governed approach can increase useful compute across diverse environments. Operators can tune the optimal operating point for each deployment based on hardware, workload mix, software, cooling, network topology, and service objectives.\n\n## Plan the site for the validated operating target\n\nDynamic allocation is an operational capability, but the site must be able to host the capacity it enables. Electrical distribution, cooling, network fabric, floor space, and rack positions should be sized for the validated lifecycle target even if fewer racks are populated on day one.\n\nFor future NVIDIA Vera Rubin NVL72 AI factories, DSX MaxLPS also combines dynamic power management with performance-per-watt techniques and infrastructure designed for 45°C liquid-cooling inlet operation. Any Vera Rubin capacity projection should remain separate from this measured GB300 NVL72 evaluation.\n\nDSX MaxLPS gives operators a framework for turning workload variability into managed capacity. The engineering work is to define the boundary, measure representative behavior, tune policy against service objectives, and prove compliance under both normal and adverse conditions.\n\nReview [NVIDIA DSX MaxLPS](https://developer.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlps/) and [NVIDIA Dynamic Power Software documentation](https://docs.nvidia.com/datacenter/dps/versions/0.8/), then use this validation sequence to establish production operating limits for your own AI factory.", "url": "https://wpnews.pro/news/how-nvidia-dsx-maxlps-maximizes-ai-factory-throughput-and-efficiency", "canonical_source": "https://developer.nvidia.com/blog/how-nvidia-dsx-maxlps-maximizes-ai-factory-throughput-and-efficiency/", "published_at": "2026-09-28 01:00:00+00:00", "updated_at": "2026-09-28 01:28:55.052538+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-chips", "large-language-models"], "entities": ["NVIDIA", "Nscale", "NVIDIA DSX MaxLPS", "Kimi K2.5", "NVIDIA GB300 NVL72", "NVIDIA Blackwell Ultra", "NVIDIA Dynamo", "NVIDIA TensorRT LLM"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-nvidia-dsx-maxlps-maximizes-ai-factory-throughput-and-efficiency", "markdown": "https://wpnews.pro/news/how-nvidia-dsx-maxlps-maximizes-ai-factory-throughput-and-efficiency.md", "text": "https://wpnews.pro/news/how-nvidia-dsx-maxlps-maximizes-ai-factory-throughput-and-efficiency.txt", "jsonld": "https://wpnews.pro/news/how-nvidia-dsx-maxlps-maximizes-ai-factory-throughput-and-efficiency.jsonld"}}