AI Infrastructure
NVIDIA Makes Performance per Watt the Vera Rubin Story
NVIDIA said on July 21 that Vera Rubin NVL72 production is ramping across a 350-plus-site supply chain and cited a CoreWeave DeepSeek-R1 benchmark reporting 10x more tokens per megawatt than Grace Blackwell NVL72. The number is a partner benchmark under a specified workload, not a universal model-speed multiplier.

What happened and why it matters
For agentic systems, the useful unit is not peak throughput alone but how much useful inference a power budget can buy.
Primary source
Primary reference: NVIDIA Blog: Vera Rubin. Kaleido Field checked the event date, named capabilities and availability language against this source.
| Source date | July 21, 2026 |
|---|---|
| Checked by Kaleido Field | July 22, 2026, 09:18 CST |
| What this source supports | NVIDIA platform announcement with a named partner benchmark for NVIDIA Vera Rubin CoreWeave 10x tokens per megawatt DeepSeek R1 July 21 2026 |
| What it does not prove | It does not prove a universal product ranking, full regional availability, or performance on every visual intelligence task. |
What changed in the announcement
NVIDIA says Vera Rubin NVL72 is ramping at partners including CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure. It describes a supply chain spanning more than 350 factory sites in 30 countries.
The company cites CoreWeave's DeepSeek-R1 benchmark on live Vera Rubin hardware: 10 times more tokens per second per megawatt than Grace Blackwell NVL72.
Why power is an agent metric
Agents add tool calls, retries, context and multi-step reasoning to the cost of a response. A platform that converts the same power budget into more completed trajectories can change the economics of long-running workloads even when the model itself is unchanged.
NVIDIA also describes codesigned CPUs, networking and cooling. That is a platform-level efficiency argument; it should be assessed at the workload and data-centre level rather than inferred from one accelerator specification.
Evidence boundary
The current figures are NVIDIA's claims, with the 10x result attributed to a CoreWeave benchmark on DeepSeek-R1. The page does not provide the complete workload configuration, prompts, quality thresholds, power measurement protocol or independent replication.
The published result should be read as a dated partner benchmark, not a general claim that every model or application will be 10 times faster or cheaper.
Evidence boundary
This page reports a dated event from a named primary source. Company specifications and adoption statements remain attributed claims unless independent evidence is cited above.
FAQ
What is the practical answer?
NVIDIA said on July 21 that Vera Rubin NVL72 production is ramping across a 350-plus-site supply chain and cited a CoreWeave DeepSeek-R1 benchmark reporting 10x more tokens per megawatt than Grace Blackwell NVL72. The number is a partner benchmark under a specified workload, not a universal model-speed multiplier.
What source does this article use?
The primary source is NVIDIA Blog: Vera Rubin. Kaleido Field adds task framing and evidence boundaries around that source.
Where should the user verify the answer?
Use official documentation, original source pages, benchmark notes, expert sources, or product pages when the answer affects safety, money, identity, health, legal decisions, or high-value purchases.