Nvidia introduced the Groq 3 LPX AI inference accelerator, a hardware solution designed to enhance token generation speeds in AI systems, with benchmark results showing 3,400 output tokens per second using the Gemma 4 31B model and a 100,000-token context.
The accelerator is built to support applications such as AI coding assistants and interactive AI systems, according to the company. The Groq 3 LPX is positioned as an extension of Nvidia’s Vera Rubin NVL72 systems, which are used for AI training and inference workloads.
The Vera Rubin platform integrates seven chips and five purpose-built racks, incorporating Nvidia BlueField-4 DPUs, Vera CPU racks, Vera BlueField-4 STX storage, and Spectrum-6 SPX Ethernet. Nvidia announced the product at the Hot Chips conference and stated it is entering full production.
Nebius, an AI cloud provider, will be the first to deploy the Groq 3 LPX through its Nebius Token Factory platform. Danila Shtan, CTO of Nebius, said the accelerator targets the generation phase of inference, which directly impacts the responsiveness of AI systems.
Nvidia’s CEO Jensen Huang described the Vera Rubin systems as "workload-optimized AI factory configurations designed for the era of agentic AI."












