NVIDIA released benchmark data showing that its Vera Rubin NVL72 systems achieve up to 30x higher throughput per megawatt compared to the GB300 NVL72 on agentic workloads. Measured using the SemiAnalysis AgentX benchmark, the test evaluated real-world coding trajectories featuring expanding context windows, sub-agent spawns, and tool calling on models such as DeepSeek V4 Pro.
Agentic AI workloads consume significantly more tokens than simple chat interfaces due to iterative reasoning steps and persistent context accumulation. NVIDIA’s DSX MaxLPS technology manages power across racks and workloads to provision up to 40% more GPUs within a single megawatt budget, helping offset the energy footprint of long-context inference.
According to the company, the efficiency gains translate directly to cost savings, offering up to 35x lower cost per million tokens than prior architectures. This allows cloud providers and enterprise data centers to run continuous multi-agent sessions at scale under tight power constraints.
Why it matters
Provides a 30x efficiency boost that enables continuous running of heavy multi-agent workflows within fixed data center power budgets.
Reduces token costs by up to 35x, drastically improving unit economics for startups building token-heavy agentic applications.
Establishes real-world multi-step agent trajectories, rather than raw single-prompt speed, as the primary benchmark for enterprise AI hardware.
Source: blogs.nvidia.com
