NVIDIA announced today that its Groq 3 LPX rack-scale system has entered full production, extending the Vera Rubin NVL72 platform to handle fast token generation for agentic systems. In an Artificial Analysis benchmark running the Gemma 4 31B model, the platform achieved 3,400 output tokens per second across 100,000-token context windows, outperforming rival platforms by 4x.
The system relies on an integrated architecture that combines compute, networking, and specialized inference acceleration. Industry partners are already deploying the platform, with Nebius becoming the first AI cloud to adopt Groq 3 LPX for its Token Factory. Additionally, CoreWeave has deployed Spectrum-X Multiplane to connect Vera Rubin racks, while SpaceXAI plans to use NVIDIA Vera CPUs for its next-generation agentic AI.
NVIDIA emphasized that as workloads transition from training to reasoning, inference demands infrastructure built specifically for throughput, low latency, and scale. By tightly integrating networking, context processing, and generation across the hardware stack, the company aims to optimize every stage of the agentic AI pipeline.
Why it matters
Enables 4x faster token generation for 100,000-token context windows, drastically reducing execution times for complex multi-agent workflows.
Nebius and CoreWeave deployments provide immediate cloud access for startups building interactive, real-time agentic applications.
Demonstrates NVIDIA’s strategy of extreme stack codesign to maintain dominance as AI workloads shift from training to reasoning.
Source: blogs.nvidia.com

