With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
17:00 · August 24, 2026 · NVIDIA

The next era of AI inference won’t be defined by a single breakthrough chip, network or system. It’ll be defined by how every layer of the AI factory works together. That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems. Announced today, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq […]
Summary
NVIDIA has placed its Vera Rubin rack-scale platform into full production, extending the NVL72 architecture with the Groq 3 LPX inference accelerator to address the token-generation demands of agentic AI workloads. The system pairs Rubin GPUs for large-context processing with LPX accelerators optimized for low-latency decode, delivering 3,400 output tokens per second on Gemma 4 31B at 100,000-token contexts—four times the rate of the nearest competing platform in Artificial Analysis benchmarks. This performance stems from tight codesign across compute, networking and software rather than isolated component improvements.
The approach targets the shift from training-centric to inference-centric AI factories, where multi-agent systems generate long token sequences, maintain extended context windows and coordinate across tools and external services. NVIDIA Spectrum-X Multiplane Ethernet scales these flows across flat, resilient fabrics that support up to 512,000 GPUs without an additional network tier, while BlueField-4 and DOCA accelerate infrastructure services such as storage access, security and observability under the new Scale-In framework. NVLink Fusion further allows custom XPUs to integrate into the same unified platform.
Early adopters include Nebius, which is deploying Groq 3 LPX as the first cloud provider to offer the accelerator to developers; CoreWeave, which has moved Spectrum-X Multiplane into production for its AI cloud; and SpaceXAI, which will use Vera CPUs for orchestration, tool use and simulation tasks across terrestrial and orbital deployments. Together these elements form a purpose-built “token factory” architecture intended to balance throughput, responsiveness and cost at the scale required by reasoning and agentic applications.
Why it matters
This article highlights a critical shift in AI infrastructure from model training to high-speed inference, which is essential for deploying responsive AI agents. For the Dutch AI market, which hosts significant European data center infrastructure, these hardware advancements dictate the future capabilities and economics of local AI services.





