Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, announced new Vera Rubin infrastructure developments and power management results at the AI Infra Summit in Santa Clara. NVIDIA reported that attendance at the event surpassed 8,000 people, up from 3,500 the previous year.
Hardware partnerships highlighted at the conference span several chip and memory architectures. Amazon’s Annapurna Labs is working with NVIDIA on custom NVHBM high-bandwidth memory technology, while d-Matrix is integrating Vera central processing units with its Raptor XPUs using NVLink Fusion. Pinterest has deployed NVIDIA’s Blackwell platform alongside Dynamo inference software for visual discovery services.
Power management results
Cloud provider Lambda published the first deployment data for NVIDIA DSX MaxLPS software on Blackwell servers. Lambda operated 19 server nodes within the power budget typically assigned to 16 nodes, increasing token throughput by 24% from four million to five million tokens per second. The setup improved performance per watt by 23%, while NVIDIA projected DSX MaxLPS can enable up to 40% more GPU capacity within the same megawatt envelope on Vera Rubin NVL72 systems.
Electricity grid integration formed another validation testing ground for the software lineup. Partner Emerald AI demonstrated automated load reduction with Silicon Valley Power, using DSX Flex within its Conductor software to adjust facility power use against utility demand signals. The software automatically reduces power to low-priority computing tasks during peak grid events and restores it once demand eases.
Benchmark performance
Independent benchmark testing showed throughput gains across multiple model types. In MLCommons MLPerf Inference v6.1 results, the Vera Rubin NVL72 system delivered up to 3.7 times higher throughput than the GB300 NVL72. A separate 288-GPU submission spanning four GB300 NVL72 racks recorded 99% scaling efficiency, while software optimizations alone drove a 1.6-times performance gain over MLPerf v6.0.
Evaluations on the SemiAnalysis AgentX benchmark recorded that Vera Rubin NVL72 produced up to 30 times higher throughput per megawatt than GB300 NVL72 on the DeepSeek V4 Pro model. NVIDIA reported the architecture achieved up to 45 times lower cost per million tokens on those tests. Combining Vera Rubin with Groq 3 LPX hardware yielded 2,529 output tokens per second per user on a 100,000-token context Qwen 3.8 27B workload.
Software startups also documented performance metrics for the standalone Vera CPU across data applications. DeepInfra recorded 2.2 times faster orchestration step latency, Perplexity measured 1.9 times faster sandbox starts on its SPACE platform, and ClickHouse listed Vera as the fastest machine tested on its ClickBench benchmark. For cluster communications, NVIDIA detailed its NVLink 6 architecture, which uses physical layer retry, forward error correction, and credit-based flow control to prevent cascading hardware stalls across multi-rack networks.
