Cloud provider Lambda achieved 24 percent higher token throughput within a fixed electricity budget using NVIDIA's DSX MaxLPS management software, NVIDIA announced on Tuesday.
The benchmark utilized a five-rack, 19-node cluster of NVIDIA HGX B200 GPU servers operating under the power limit normally allocated to 16 nodes at full capacity. Output increased from roughly 4 million tokens per second to 5 million, raising performance per watt by 23 percent. NVIDIA projects the same software can support up to 40 percent more GPU capacity in next-generation Vera Rubin NVL72 facilities without increasing megawatt allocations.
Dave Ward, president of cloud services at Lambda, said the software reclaims stranded capacity by monitoring GPU and rack-level electricity consumption. MaxLPS dynamically shifts headroom between inference and training workloads, which draw power at differing rates, rather than relying on static rack provisioning.
Automated grid coordination trials in Santa Clara, California, demonstrated complementary load-reduction mechanisms. NVIDIA's Eos AI factory responded to grid stress signals sent by municipal utility Silicon Valley Power under its Flexible Load Interconnect Program. Using Emerald AI's Conductor platform, the data center cut load from four megawatts to three megawatts in under a minute by pausing low-priority jobs while maintaining high-priority inference tasks.
Silicon Valley Power has issued more than 200 successful demand signals to the Eos facility without dropped workloads, according to NVIDIA. The company plans to integrate Conductor into its DSX Flex platform ahead of a commercial rollout at a 96-megawatt Vera Rubin site in Manassas, Virginia. NVIDIA is also incorporating an 800-volt direct-current architecture into its hardware reference designs to reduce power conversion losses across dense computing racks.
