HomeTechHardwareNVIDIA Puts Groq 3 LPX Into Full Produ
HARDWARE

NVIDIA Puts Groq 3 LPX Into Full Production

NVIDIA launched Groq 3 LPX into production alongside new networking and custom silicon architectures for agentic AI workloads.

WHAT YOU NEED TO KNOW
  • Groq 3 LPX delivered 3,400 output tokens per second running Gemma 4 31B over 100,000-token context windows.
  • Nebius became the first AI cloud provider to deploy Groq 3 LPX alongside Vera Rubin NVL72 racks.
  • SpaceXAI will deploy Vera CPUs for orchestration and simulation across data centers and satellites.
  • Spectrum-X Multiplane scales two-tier networks up to 512,000 GPUs using ConnectX-9 SuperNICs and SN6000 switches.

NVIDIA announced that its rack-scale Groq 3 LPX system has entered full production, extending the Vera Rubin NVL72 platform for agentic AI inference workloads. In an Artificial Analysis benchmark running the Gemma 4 31B model, the system generated 3,400 output tokens per second across 100,000-token context windows, which NVIDIA said is four times faster than the nearest alternative platform.

Cloud and enterprise providers have started adopting the hardware. Nebius is the first AI cloud provider to adopt Groq 3 LPX, integrating the accelerators into its Nebius Token Factory alongside Vera Rubin NVL72 systems. SpaceXAI plans to deploy NVIDIA Vera CPUs to run orchestration, tool use, code execution, data processing, and simulation across terrestrial data centers and orbital satellites. CoreWeave deployed NVIDIA Spectrum-X Multiplane into production to link its Vera Rubin racks.

Accelerator architecture

The Groq 3 LPX architecture divides inference tasks across specialized silicon. Rubin GPUs handle context processing, while LPX accelerators manage token decode workloads. At rack scale, a Groq 3 LPX deployment connects up to 256 LP30 accelerators through direct chip-to-chip links.

Networking and custom silicon

At the Hot Chips conference in Palo Alto, California, NVIDIA showcased Spectrum-X Multiplane networking. The architecture splits server connections across parallel planes running lightweight two-tier networks, scaling up to 512,000 GPUs without a third network tier. A hardware engine inside the ConnectX SuperNIC manages plane routing. In an eight-plane topology, a single failed plane leaves about 90% of total bandwidth intact, with hardware recovery operating 11 times faster than software-based load balancing.

The networking platform relies on Spectrum-X SN6000 switches using the 102.4Tb/s Spectrum-6 ASIC and ConnectX-9 SuperNICs supporting up to 1,600Gb/s per GPU. Across separate data centers, Spectrum-XGS Ethernet accelerates multi-site NCCL collectives by 1.9 times.

NVIDIA also introduced Scale-In and NVLink Fusion. Scale-In uses BlueField-4 processors and the DOCA software platform over Spectrum-X Ethernet to accelerate storage access, multi-tenant networking, in-silicon security, and provisioning without drawing on host compute resources. NVLink Fusion connects third-party XPUs and CPUs to sixth-generation NVLink, NVLink Switch, and NVLink-C2C interconnects within the NVIDIA MGX rack ecosystem.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →