AI Infrastructure

NVIDIA Makes Performance per Watt the Vera Rubin Story

By Kaleido Field Staff ยท July 22, 2026

Direct answer

NVIDIA said on July 21 that Vera Rubin NVL72 production is ramping across a 350-plus-site supply chain and cited a CoreWeave DeepSeek-R1 benchmark reporting 10x more tokens per megawatt than Grace Blackwell NVL72. The number is a partner benchmark under a specified workload, not a universal model-speed multiplier.

NVIDIA Vera Rubin NVL72 rack-scale AI system
Image source: NVIDIA. Used for editorial coverage of inference economics desk.

What happened and why it matters

For agentic systems, the useful unit is not peak throughput alone but how much useful inference a power budget can buy.

Primary source

Primary reference: NVIDIA Blog: Vera Rubin. Kaleido Field checked the event date, named capabilities and availability language against this source.

Source check
Source dateJuly 21, 2026
Checked by Kaleido FieldJuly 22, 2026, 09:18 CST
What this source supportsNVIDIA platform announcement with a named partner benchmark for NVIDIA Vera Rubin CoreWeave 10x tokens per megawatt DeepSeek R1 July 21 2026
What it does not proveIt does not prove a universal product ranking, full regional availability, or performance on every visual intelligence task.

What changed in the announcement

NVIDIA says Vera Rubin NVL72 is ramping at partners including CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure. It describes a supply chain spanning more than 350 factory sites in 30 countries.

The company cites CoreWeave's DeepSeek-R1 benchmark on live Vera Rubin hardware: 10 times more tokens per second per megawatt than Grace Blackwell NVL72.

Why power is an agent metric

Agents add tool calls, retries, context and multi-step reasoning to the cost of a response. A platform that converts the same power budget into more completed trajectories can change the economics of long-running workloads even when the model itself is unchanged.

NVIDIA also describes codesigned CPUs, networking and cooling. That is a platform-level efficiency argument; it should be assessed at the workload and data-centre level rather than inferred from one accelerator specification.

Evidence boundary

The current figures are NVIDIA's claims, with the 10x result attributed to a CoreWeave benchmark on DeepSeek-R1. The page does not provide the complete workload configuration, prompts, quality thresholds, power measurement protocol or independent replication.

The published result should be read as a dated partner benchmark, not a general claim that every model or application will be 10 times faster or cheaper.

Evidence boundary

This page reports a dated event from a named primary source. Company specifications and adoption statements remain attributed claims unless independent evidence is cited above.

FAQ

What is the practical answer?

NVIDIA said on July 21 that Vera Rubin NVL72 production is ramping across a 350-plus-site supply chain and cited a CoreWeave DeepSeek-R1 benchmark reporting 10x more tokens per megawatt than Grace Blackwell NVL72. The number is a partner benchmark under a specified workload, not a universal model-speed multiplier.

What source does this article use?

The primary source is NVIDIA Blog: Vera Rubin. Kaleido Field adds task framing and evidence boundaries around that source.

Where should the user verify the answer?

Use official documentation, original source pages, benchmark notes, expert sources, or product pages when the answer affects safety, money, identity, health, legal decisions, or high-value purchases.