bytevyte
bytevyte
Language
ai-beats

Nvidia Puts Tokens per Megawatt at the Centre of the AI Factory Pitch

tokens per megawatt

Nvidia has recast the economics of AI data centres around a single metric, tokens per megawatt, and used its AI Infra Summit in Santa Clara to attach hard multipliers to it. The company says the Vera Rubin NVL72 rack-scale system delivers up to 30x higher throughput per megawatt than earlier generations, up to 35x higher token throughput per megawatt on large models, and up to 45x lower cost per million tokens on agentic workloads. Vera Rubin is in full production, Nvidia stated. Summit attendance reached more than 8,000 people, up from 3,500 a year earlier.

The framing carries as much weight as the figures. Nvidia's argument is that an AI factory should be judged the way a power plant is judged, on useful output per unit of energy consumed, rather than on accelerator specifications alone. Ian Buck, Nvidia's vice president of hyperscale and HPC, presented Vera Rubin and the DSX platform as one system spanning silicon, power delivery and grid interaction.

The 30x Tokens per Megawatt Claim

The headline numbers come from agentic coding workloads. Nvidia measured on-silicon data using the SemiAnalysis AgentX benchmark, with results quoted at 160 tokens per second per user on a DeepSeek V4-Pro workload. Against that yardstick, the company claims up to 30x higher AI factory throughput per megawatt and up to 35x lower token cost than its previous flagship system.

One detail complicates the comparison. Nvidia's own materials state the baseline inconsistently: some describe the 30x gain against GB300 NVL72, the current shipping flagship, while others cite GB200 NVL72, the generation before it. The distinction is not cosmetic. A 30x multiple measured against an older part describes a different upgrade case, and therefore a different replacement cycle, than the same number measured against the hardware customers bought most recently.

MetricNvidia claimBaseline
AI factory throughput per MWup to 30x higherGB300 NVL72 (also cited as GB200 NVL72)
Token throughput per MW, large modelsup to 35x higherGB200 NVL72
Cost per million tokensup to 45x lowerAgentX agentic benchmarks
GPU capacity in a fixed power budgetup to 40% moreDSX MaxLPS, suitable environments
Token throughput, same power envelope24% moreLambda validation of MaxLPS
Performance per watt23% improvementNvidia power management stack

Nvidia also reported a validation from Lambda, which found that DSX MaxLPS supports 24% more token throughput inside an unchanged power envelope, and a 23% improvement in performance per watt from the company's power management stack. CoreWeave's first benchmark on DeepSeek-R1 showed 10x more throughput per megawatt than Grace Blackwell NVL72, a smaller multiple that reflects a different model and a customer-run test rather than a vendor-run one.

Nvidia frames the metric commercially as well as technically. Its developer material cites a 10x expansion in revenue per gigawatt for Vera Rubin paired with Groq 3 LPX, which turns tokens per megawatt into an argument about asset returns rather than engineering elegance. That framing suits operators who finance capacity against expected output, because it lets them compare an AI factory with any other energy-consuming asset on the revenue it produces per unit of grid capacity.

From Power Buyer to Grid Asset

The second half of Nvidia's pitch concerns interconnection. DSX Flex lets an AI factory respond to utility demand signals, cutting consumption when the grid is strained and restoring it afterwards. In a pilot with Silicon Valley Power, an AI factory automatically reduced load from 4MW to 3MW without dropping high-priority jobs, with the coordination handled through Emerald AI.

Nvidia's first dedicated DSX Flex commercial deployment is planned for a facility in Manassas, Virginia. The company has lined up energy partners including AES, Constellation, Invenergy, NextEra Energy, Nscale Energy & Power and Vistra, a group it announced alongside Emerald AI at CERAWeek in March 2026 with the aim of connecting AI factories to the grid faster while letting them operate as flexible energy assets.

Hardware roadmaps are shifting to match. Future Nvidia reference designs will move to an 800V DC power architecture, which reduces the number of conversion stages between grid supply and accelerator. The Omniverse DSX Blueprint is the operating layer Nvidia puts around that hardware to hold a facility at peak efficiency, alongside power smoothing, local energy buffering and grid-aware control.

Inside the DSX Toolchain

Nvidia splits the DSX software stack by job. MaxLPS allocates power across GPUs, racks and workloads. Max-Q targets the maximum number of AI tokens per watt of available energy. Flex handles utility demand response. The underlying Vera Rubin DSX reference design is open, modular and composable, so operators can combine components without rebuilding an entire system, and Nvidia presents it as a guide that shortens time to first production.

Nvidia released that reference design alongside an Omniverse DSX digital twin blueprint, both aimed at maximum tokens per watt and faster time to first production. The digital twin approach lets operators model power, cooling and workload behaviour before a facility is built, which matters when interconnection timelines, rather than chip supply, determine when capacity goes live.

The packaging matters as much as the software. Nvidia bundles CPU, GPU, networking and storage into pod-scale units designed, tested and operated as a single system rather than as separate parts. Vera, described as the CPU for agents, handles the CPU-heavy side of that unit: code execution, environment stepping, data preparation and control logic. Full-stack confidential computing extends trust across CPUs, GPUs and interconnects at rack scale.

Vera Rubin's design assumption is that agentic AI changes the ratio of work between CPU and GPU. An agent that plans, calls tools, executes code and checks results spends far more time in orchestration than a single chat completion does, which is why Nvidia pairs Rubin GPUs with the Vera CPU and reports analytical query and orchestration latency gains for users including Perplexity and ClickHouse.

The Trade-offs Behind the Multipliers

Two caveats qualify the efficiency story. The first is measurement. The 30x, 35x and 45x figures come from Nvidia's own testing on selected workloads, and the gains narrow considerably in data produced by customers, as the CoreWeave result shows. Tokens per megawatt also varies with model, batch size, sequence length and the latency target imposed on each user, which makes cross-vendor comparisons unreliable without a shared workload specification.

The second is the difference between capacity claims. Nvidia says DSX MaxLPS can add up to 40% more GPU capacity inside a fixed power budget in suitable environments, while platform-level features such as power smoothing and energy buffering are credited with up to 30% more GPU capacity per megawatt. Both depend on site conditions, including cooling design, grid quality and workload mix. Not every facility can realise them.

A demand-side effect also deserves attention. Cheaper tokens lower the cost of running inference, which historically raises total consumption rather than reducing it, because more workloads become economically viable at a lower price per million tokens. Nvidia's partner pipeline shows the scale involved: a Noetra Corp. buildout described as the first national AI infrastructure project covers 13,750 Vera CPUs and 27,500 Rubin GPUs across 140 megawatts of data centre capacity.

Why this matters

Nvidia is selling a metric that reframes its customers' binding constraint. Power, not silicon, limits AI capacity in most markets, so a vendor promising more useful output per megawatt is competing on the budget line that actually blocks expansion. That favours operators with secured interconnection and flexible loads, and it resets the comparison against older installed fleets whose efficiency per megawatt is now the yardstick.

For buyers, the practical test is replication rather than the summit figures. The Manassas DSX Flex deployment, the 800V DC transition and independent benchmarks on production workloads will show how much of the 30x survives contact with a real facility.

Sources

AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories

From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production

Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

NVIDIA Corporation - NVIDIA and Emerald AI Join Leading Energy Companies to Pioneer Flexible AI Factories as Grid Assets

Scaling Token Factory Revenue and AI Efficiency by Maximizing Performance per Watt | NVIDIA Technical Blog

NVIDIA Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance per Watt | NVIDIA Technical Blog

NVIDIA Vera Rubin Driving Performance Per Watt, Lower Token Costs for Partners Worldwide | NVIDIA Blog

Inside the NVIDIA Vera Rubin Platform: Six New Chips, One AI Supercomputer | NVIDIA Technical Blog

NVIDIA Corporation - NVIDIA Releases Vera Rubin DSX AI Factory Reference Design and Omniverse DSX Digital Twin Blueprint With Broad Industry Support

Photo by Brecht Corbeel on Unsplash

✔Human Verified


Researched and cross-referenced against primary sources by the Bytevyte editorial team. This article was generated with the assistance of artificial intelligence and reviewed by the Bytevyte editorial team.