Actualité

Nvidia’s New Power-Management Software Squeezes Up to 40% More GPUs Into the Same Energy Budget

Nvidia says its DSX MaxLPS software can pack more GPUs into existing power envelopes, with early real-world tests from cloud provider Lambda backing up the pitch. The company frames electricity, not raw chip speed, as the new bottleneck for AI infrastructure.

Nvidia is making the case that the real constraint on building AI infrastructure isn’t chip performance anymore — it’s electricity. At the AI Infra Summit on September 15, 2026, the company unveiled DSX MaxLPS, software it says can let data centers cram up to 40% more GPUs into the same power budget when running Vera Rubin NVL72 “AI factories,” assuming conditions are favorable.

The claim isn’t purely theoretical. Cloud provider Lambda ran an early test on Blackwell-based servers and got a more modest but tangible result: 19 nodes operated within a power budget normally sized for 16 nodes running flat out. Nvidia says that translated into overall throughput climbing from roughly 4 million to 5 million tokens per second, a 24% jump, along with a 23% gain in performance per watt.

The software works by watching power draw at both the GPU and rack level, then shifting available capacity around in real time. The logic behind it is simple: not every piece of hardware peaks at the same moment, and training workloads pull power very differently than inference does. That leaves slack in systems that are normally provisioned statically for worst-case demand — slack DSX MaxLPS is designed to claw back.

Nvidia’s broader numbers are bigger, and it’s careful to call them ceilings rather than guarantees. Site-wide optimization, the company says, could deliver up to 1.4 times more tokens per megawatt, and a Vera Rubin NVL72 deployment could see up to 35% more throughput without adding any new power lines. How close any given data center gets to those figures depends heavily on its mix of workloads, how variable they are, and what electrical infrastructure is already in place.

Behind the push is the growing appetite of agentic AI workloads, which chain together multiple reasoning steps, tool calls, and long context windows — all of which multiply the compute needed per session. That’s why Nvidia is increasingly talking in terms of tokens per megawatt rather than raw chip speed.

The company also shared fresh Vera Rubin NVL72 benchmarks using AgentX, a test built by SemiAnalysis around recorded agentic coding sessions. Running DeepSeek V4 Pro, Vera Rubin reportedly hit up to 30 times the throughput per megawatt of the previous-generation GB300 NVL72, while cutting the cost per million tokens by as much as 45 times. AgentX tries to mimic how a real session unfolds — growing context, delays from tool calls, sub-agents spinning up — and Nvidia says such sessions can process about 15 times more tokens than a simple chat query, so these multipliers shouldn’t be assumed to carry over to every model or use case.

Those gains lean on a stack of technology working together: the NVL72 compute domain, sixth-generation NVLink, NVFP4 precision on fifth-generation Tensor Cores, plus TensorRT LLM and Dynamo software. Nvidia says Vera Rubin is already shipping at scale across its ecosystem, with more software tuning still to come.

For low-latency inference, Nvidia’s Groq 3 LPX is being pitched as a complementary option, claiming up to 35 times more tokens per megawatt than GB200 NVL72 for large, long-context models above 2 trillion parameters. In one test using Qwen 3.8 27B with a 100,000-token context, it produced 2,529 output tokens per second per user — again, a result tied to a specific setup.

Nvidia also showed off a demand-flexibility experiment with Emerald AI and utility Silicon Valley Power, where a data center responded to hundreds of grid signals while shielding its highest-priority AI jobs. Emerald AI plans to fold this into its Conductor software using a new Nvidia tool called DSX Flex, which can throttle or pause lower-priority tasks in response to utility signals like demand response events or pricing changes, then resume them once the pressure eases — effectively letting a data center act as a flexible load on the grid.

Finally, Nvidia detailed resilience features built into NVLink 6 aimed at systems with potentially hundreds of thousands of GPUs, combining error correction, retransmission, and link recovery at the physical layer with dynamic routing and load balancing at the network level, all meant to catch failures before they ripple out. Taken together, Nvidia is positioning Vera Rubin, NVLink, Dynamo, BlueField DPUs, and its DSX software suite as one integrated system — from the silicon to the power grid connection — even as it acknowledges its flashiest numbers remain closely tied to the specific models, software, and test conditions used to produce them.

Leave a comment

Your email address will not be published.