NVIDIA has released performance benchmarks for its upcoming Vera Rubin platform that signal a major shift in how data centers will handle large-scale AI tasks. These numbers matter because they suggest cloud providers can run significantly more AI agents for less money than current hardware allows. The Vera Rubin NVL72 system delivers a 30x increase in throughput per watt compared to the Grace Blackwell NVL72 in DeepSeek-v4-PRO 1.6T workloads. This efficiency gain means operators can squeeze more compute power from the same electrical budget.

New benchmarks show massive efficiency gains for agentic AI workloads
The Vera Rubin architecture targets Agentic AI workloads where systems must perform complex, multi-step coding and reasoning tasks. NVIDIA claims the platform offers a 35x reduction in token cost per million tokens compared to Blackwell for these specific workloads. Lower costs per token allow AI factories to run several agents continuously at scale across a broad set of applications. This pricing structure fundamentally changes the economics of deploying autonomous AI software at the enterprise level.
- Throughput Increase vs Blackwell: 30x
- Token Cost Reduction vs Blackwell: 35x
- Interactivity (TPS): ~160-280 tokens per second per user
- GPU Provisioning Increase: 40% more GPUs within the same Megawatt budget
Technical benchmarks show the Vera Rubin system achieves interactivity rates of approximately 160 to 280 tokens per second per user. This speed allows the system to respond to user inputs much faster than the Grace Blackwell architecture. AI factories can provision up to 40% more GPUs within the same megawatt budget using the new platform. The increased density helps operators maximize their infrastructure investment without expanding their physical footprint or power consumption.
Previous generations like the H200 NVL8 offered 15x higher throughput per megawatt than older Hopper solutions in similar DeepSeek workloads. Blackwell itself improved upon Hopper, but Vera Rubin pushes the efficiency curve even further. NVIDIA has commenced production across Vera CPUs, Rubin GPUs, and Vera Rubin servers to support these new capabilities. The company notes that current benchmarks focus on the Vera Rubin chips rather than the complete seven-chip Extreme Codesign platform.
These benchmarks do not yet reflect Vera CPU performance for tool calling, which remains a key part of the full system architecture. The reported figures isolate the GPU and specific workload optimizations to highlight the raw compute gains. NVIDIA positions Vera Rubin as a platform that offers disruptive token throughput at a much lower cost than Blackwell. Data centers planning their next AI infrastructure upgrades should monitor these efficiency metrics closely as they finalize their procurement strategies.



Discussion
0 comments
Log in to join the thread with a thoughtful take, question, or correction.