NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut
NVIDIA's Vera Rubin NVL72 system debuts with leading performance in MLPerf Inference v6.1, delivering up to 3.7x better throughput than GB300 NVL72.
NVIDIA announced the debut of its Vera Rubin NVL72 system in the MLPerf Inference v6.1 benchmark, showcasing leading performance in AI inference. In its first preview submission, the Vera Rubin NVL72 system delivered up to 3.7x better throughput than the GB300 NVL72 system. The system achieved these results on two of the most demanding benchmarks in the MLPerf Inference v6.1 suite: DeepSeek-R1 and Qwen3-VL. On Qwen3-VL, Vera Rubin NVL72 demonstrated up to 3.7x higher throughput across offline, server, and interactive scenarios, while on DeepSeek-R1, it showed up to 2.5x higher throughput. These early results highlight NVIDIA's accelerated pace of innovation and the potential for further performance gains through continuous software optimizations. The NVIDIA platform is designed to optimize system performance, efficient infrastructure scaling, and continuous software optimization, key factors in determining AI inference economics. Each Vera Rubin NVL72 rack delivers significantly more tokens, serves more users, and generates more revenue than a GB300 NVL72 rack, while also lowering the cost per token.
Source: nvidia-newsroom



