NVIDIA announced that its rack-scale Groq 3 LPX system has entered full production, extending the Vera Rubin NVL72 platform for agentic AI inference workloads. In an Artificial Analysis benchmark running the Gemma 4 31B model, the system generated 3,400 output tokens per second across 100,000-token context windows, which NVIDIA said is four times faster than the nearest alternative platform.
Cloud and enterprise providers have started adopting the hardware. Nebius is the first AI cloud provider to adopt Groq 3 LPX, integrating the accelerators into its Nebius Token Factory alongside Vera Rubin NVL72 systems. SpaceXAI plans to deploy NVIDIA Vera CPUs to run orchestration, tool use, code execution, data processing, and simulation across terrestrial data centers and orbital satellites. CoreWeave deployed NVIDIA Spectrum-X Multiplane into production to link its Vera Rubin racks.
Accelerator architecture
The Groq 3 LPX architecture divides inference tasks across specialized silicon. Rubin GPUs handle context processing, while LPX accelerators manage token decode workloads. At rack scale, a Groq 3 LPX deployment connects up to 256 LP30 accelerators through direct chip-to-chip links.
Networking and custom silicon
At the Hot Chips conference in Palo Alto, California, NVIDIA showcased Spectrum-X Multiplane networking. The architecture splits server connections across parallel planes running lightweight two-tier networks, scaling up to 512,000 GPUs without a third network tier. A hardware engine inside the ConnectX SuperNIC manages plane routing. In an eight-plane topology, a single failed plane leaves about 90% of total bandwidth intact, with hardware recovery operating 11 times faster than software-based load balancing.
The networking platform relies on Spectrum-X SN6000 switches using the 102.4Tb/s Spectrum-6 ASIC and ConnectX-9 SuperNICs supporting up to 1,600Gb/s per GPU. Across separate data centers, Spectrum-XGS Ethernet accelerates multi-site NCCL collectives by 1.9 times.
NVIDIA also introduced Scale-In and NVLink Fusion. Scale-In uses BlueField-4 processors and the DOCA software platform over Spectrum-X Ethernet to accelerate storage access, multi-tenant networking, in-silicon security, and provisioning without drawing on host compute resources. NVLink Fusion connects third-party XPUs and CPUs to sixth-generation NVLink, NVLink Switch, and NVLink-C2C interconnects within the NVIDIA MGX rack ecosystem.
