NVIDIA has submitted initial MLPerf Inference v6.1 performance benchmark results for its Vera Rubin NVL72 system, demonstrating significant gains in AI inference throughput. Operating on the Qwen3-VL benchmark, the Vera Rubin NVL72 delivered up to 3.7x higher throughput than the previous-generation GB300 NVL72 using the vLLM and NVIDIA Dynamo frameworks. On the DeepSeek-R1 benchmark using NVIDIA TensorRT-LLM, throughput increased up to 2.5x compared to the GB300 NVL72.
The throughput performance gains stem from hardware and software co-design, including enhanced Tensor Cores, NVFP4 precision, and disaggregated serving that separates prefill and decode stages. The NVL72 system relies on sixth-generation NVLink and NVLink Switch technology, providing 10x higher packet rates and 3x lower latency compared to standard Ethernet. Partner Nebius also submitted preview results confirming the platform’s performance.
In early testing for agentic workloads using the SemiAnalysis AgentX benchmark, the Vera Rubin NVL72 posted 30x higher performance than the GB300 NVL72. These efficiency gains aim to lower the cost per token and maximize token generation per rack for enterprise inference deployments.
Why it matters
Significantly higher token throughput per rack reduces overall unit economics and operational expenditure for large-scale enterprise AI inference.
Hardware-software co-design using NVFP4 precision and disaggregated serving sets new benchmarks for serving complex MoE reasoning models.
30x performance boosts on agentic benchmarks reflect hardware optimization tailored specifically for multi-step reasoning and autonomous AI agents.
Source: blogs.nvidia.com



