NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide
17:36 · July 21, 2026 · NVIDIA

NVIDIA Vera Rubin is here, and it’s going gigascale. Vera Rubin NVL72 production is ramping up with racks running at partners CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure. Spanning 350+ factory sites in 30 countries, Vera Rubin has the largest, most mature rack-scale supply chain ever assembled to meet customer compute demand. The […]
Summary
NVIDIA’s Vera Rubin NVL72 rack-scale system is now entering volume production and is already running at CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure. The platform integrates seven new chips and five rack trays that were codesigned as a single system, including the Vera CPU, NVLink 6 fabric, Spectrum-6 Ethernet switches and BlueField-4 data-processing units. Early benchmarks from CoreWeave on DeepSeek-R1 show a measured 10× gain in tokens per second per megawatt relative to Grace Blackwell NVL72, the metric that directly governs power-constrained AI-factory economics.
The same architecture supports agentic workloads that can consume up to 15× more tokens than conventional inference. The Vera CPU delivers 2× single-threaded performance and 3× core-to-core bandwidth over competing designs, while the 260 TB/s NVLink 6 fabric removes all-to-all communication bottlenecks that arise in mixture-of-experts models. Liquid cooling rated for 45 °C inlet temperature eliminates chillers and is projected to save millions of gallons of water per megawatt annually.
In Europe, NVIDIA, Microsoft and Mistral have announced a multibillion-dollar expansion of their partnership that will deploy tens of thousands of Vera Rubin GPUs across sovereign cloud and on-premises environments. Mistral’s open models, including Medium 3.5, will run in Microsoft Azure, Azure Local and Foundry Local instances, giving regulated sectors the option to keep data and governance controls inside the region while retaining access to frontier-scale training and inference capacity.
Why it matters
This article is highly relevant as it details a major leap in AI hardware efficiency and a massive investment in European sovereign AI infrastructure. For the Dutch market, this means access to powerful, energy-efficient AI compute that strictly adheres to EU data and governance regulations.







