SemiAnalysis has published what it describes as the first verified agentic inference results for NVIDIA's Vera Rubin NVL72 platform, measured using its AgentX benchmark across a fleet of thousands of chips. At 170 tokens per second, Vera Rubin NVL72 delivered approximately 67x the total throughput per TCO compared to the GB300 Dynamo configuration under owning-cost assumptions, and achieved up to 7x better token throughput per megawatt on pre-release software. The benchmark has been validated by major compute buyers including Google Cloud, Microsoft Azure, Oracle, and Meta, and supported by frameworks including vLLM, SGLang, and PyTorch. SemiAnalysis also estimates that even on early software builds, Rubin can earn over 2x more profit per gigawatt than the Blackwell platform, with that gap expected to widen as the software stack matures.

Why this matters

The scale of the reported efficiency gains, 67x throughput per dollar over GB300 at a specific operating point, directly affects capital allocation decisions for inference providers and hyperscale AI labs evaluating which hardware generation to deploy next. If the results hold as Rubin's software stack matures, they would establish a new cost baseline for large-scale agentic inference and accelerate the obsolescence of current Blackwell deployments.

Why the Digest selected this story

SemiAnalysis is a high-authority source on AI compute infrastructure, and a 67x performance-per-dollar claim for NVIDIA's Vera Rubin NVL72 platform is a significant benchmark finding with direct implications for hyperscaler GPU procurement and data center design decisions.

Read the full story at SemiAnalysis →