Fastest, Largest, Strongest: NVIDIA Blackwell Sweeps MLPerf Training 6.0
17:00 · June 16, 2026 · NVIDIA

Every breakthrough AI model starts the same way: with a training run. The infrastructure running those training jobs shapes everything: how fast teams can iterate, what scale of model they can build and whether those jobs complete reliably. As models grow in size, complexity and intelligence, the demands on training infrastructure are also rising. In […]
Summary
NVIDIA’s Blackwell platform achieved the fastest training times across all seven workloads in the MLPerf Training 6.0 benchmark suite, including two newly added mixture-of-experts pretraining tasks: DeepSeek-V3 671B and GPT-OSS-20B. The company submitted results on every benchmark, using both GB200 NVL72 and GB300 NVL72 rack-scale systems. Within each rack, fifth-generation NVLink connects 72 GPUs into a single high-bandwidth domain that functions as one large accelerator, which proves especially effective for the all-to-all token routing required by large MoE models.
The GB300 NVL72 variant delivered up to 1.6 times the performance of the GB200 NVL72 at equivalent scale, driven by higher compute density through NVFP4 low-precision arithmetic, increased memory capacity, and a higher sustained power limit. NVIDIA scaled its largest submission to 8,192 GPUs on DeepSeek-V3 671B and to 5,120 GPUs on the dense Llama 3.1 405B model, marking the largest Blackwell-based clusters reported in this round. Complementary scale-out fabrics—Quantum InfiniBand and Spectrum-X Ethernet—support these distributed runs while NVFP4 methods maintain accuracy across pretraining and fine-tuning workloads.
Production-scale training also depends on system resiliency over weeks or months of operation. The submitted results reflect co-engineering across hardware, networking, and software that improves both raw throughput and job reproducibility. Nineteen partner organizations, including CoreWeave, Google Cloud, and Nebius, contributed additional Blackwell-based entries, with reported gains such as threefold faster training for Cohere’s agentic platform and a 30 percent reduction in training time for Higgsfield’s models serving millions of users. These outcomes illustrate how the same platform characteristics translate into measurable iteration speed and cost reductions for frontier model development.
Why it matters
This article details crucial advancements in AI hardware infrastructure that dictate the speed and scale of future AI models. However, its heavy reliance on technical jargon makes it less accessible for a general audience, though the underlying trend impacts the entire AI ecosystem.





