Back to the stories

NVIDIA claims Vera Rubin NVL72 delivers over 30x throughput per megawatt versus GB300 NVL72 for inference

Score 9.2,

AI Infrastructure

Running large AI models is suddenly far cheaper and far more power-efficient for data centers.

NVIDIA is calling a new approach "AI factories", facilities that treat GPU fleets as durable, fungible assets. It published data showing its Vera Rubin NVL72 systems, combined with CUDA-X optimizations, outperform older GB300 NVL72 rigs when running the DeepSeek V4 Pro model.

Why that matters: power and hardware costs dominate large-scale inference. If a data center can squeeze many more tokens from the same power draw, the unit economics of serving LLMs and vision models change.

How it works, in plain terms: think of a factory line. Instead of building a machine for one job, NVIDIA tunes hardware, software and operations so each GPU does useful work longer and more efficiently, and the software stack steers workloads to the best resources.

NVIDIA cites SemiAnalysis AgentX data claiming more than 30x higher throughput per megawatt and up to 45x lower cost per million tokens for Vera Rubin NVL72 versus GB300 NVL72 on DeepSeek V4 Pro. If those numbers hold up across other models, running inference at scale becomes dramatically cheaper.

A reality check: these gains show up at data-center scale. They don't mean you can run these large models on a laptop, you still need enterprise-grade GPU infrastructure.

The open question is simple: will independent benchmarks and live deployments reproduce these gains across different models and cloud environments?