How NVIDIA Frames the Return Math for AI Factories

NVIDIA is arguing that AI factory returns hinge on three variables: earning capacity, useful life and demand breadth, and that strength in one cannot fully offset weakness in another, according to a company blog post published October 1, 2026. The post cites roughly $60 million per megawatt of capacity and SemiAnalysis AgentX data claiming Vera Rubin NVL72 delivers over 30x higher throughput per megawatt than GB300 NVL72. It also argues older GPUs keep earning, pointing to the A100 remaining in commercial service six years after its 2020 ship date.

Published: October 2, 2026 By Marcus Rodriguez, Robotics & AI Systems Editor AI Author Category: AI Chips

Marcus specializes in robotics, life sciences, conversational AI, agentic systems, climate tech, fintech automation, and aerospace innovation. Expert in AI systems and automation

How NVIDIA Frames the Return Math for AI Factories

Executive Summary

  • NVIDIA is framing the economics of AI data centers around three measurable variables: earning capacity, useful life and demand breadth, arguing that strength in one cannot fully offset weakness in another, according to a company blog post published October 1, 2026 by Shruti Koparkar.
  • Each megawatt of AI factory capacity costs roughly $60 million, a capital scale that makes return on investment the deciding factor for operators, the post states.
  • NVIDIA claims its Vera Rubin NVL72 systems deliver over 30x higher throughput per megawatt than GB300 NVL72 and up to 45x lower cost per million tokens on the DeepSeek V4 Pro model, citing SemiAnalysis AgentX data in the same post.
  • The post argues older hardware keeps earning: the A100 GPU shipped in 2020 and remains in commercial service six years later, with CoreWeave extending bookings for 2020-introduced units through 2029, per the company's account.
  • NVIDIA positions CUDA and more than 1,000 CUDA-X libraries as the mechanism that lets one architecture run AI and non-AI workloads, supported by named customer examples including Lilly, Pinterest, Revolut, Runway and Texas A&M University.

Key Takeaways

  • NVIDIA's return-on-investment case rests on tokens per second per megawatt and cost per million tokens, not on peak benchmark figures alone.
  • The company argues cheaper tokens expand rather than shrink compute demand, because newly economical use cases consume more capacity than efficiency frees up.
  • Durability claims are supported by third-party resale and rental data points the post cites, though those figures come from the vendors named in the article rather than independent audits.
  • A planned GTC Berlin keynote by founder and CEO Jensen Huang on Wednesday, Oct. 21, at 11 a.m. CEST is the next scheduled venue for NVIDIA to expand on this positioning.

Why NVIDIA Is Framing AI Factories Around Three Variables

The blog post's central argument is that AI factories are capital projects judged by return, and that three forces determine whether that return materializes. Earning capacity is what a factory could collect in a year if it sold every token it can produce. Useful life is how long its AI hardware keeps generating revenue. Demand is how much appetite exists for those tokens.

NVIDIA states that strength in one variable cannot fully offset weakness in another. High earning capacity counts for little if a factory sells only part of its output. High demand matters little if production falls to partial capacity within a year. The three also interact: a factory that can run more kinds of workloads finds more demand and keeps earning longer.

That framing is doing commercial work. Operators weighing roughly $60 million per megawatt commitments need a defensible model for depreciation, utilization and workload mix. NVIDIA's answer is that its platform is designed to maximize all three variables simultaneously, and the remainder of the post is organized as evidence for that claim.

Productive How Vera Rubin NVL72 Changes the Token Math

Because power is the binding constraint on an AI factory, NVIDIA argues tokens per second per megawatt is the number that governs earning capacity. More tokens inside a fixed power envelope means more revenue; lower cost per token means more margin on each one.

The post cites SemiAnalysis AgentX data showing NVIDIA Vera Rubin NVL72 systems deliver over 30x higher throughput per megawatt than NVIDIA GB300 NVL72, and up to 45x lower cost per million tokens on the DeepSeek V4 Pro model. It attributes gains of that size to codesign across models, workloads, software, compute, networking and memory optimized together. A separate SemiAnalysis AgentX analysis cited in the post concerns GB300 NVL72 performance for GLM 5.3.

NVIDIA then addresses the obvious follow-on question. If every generation makes tokens dramatically cheaper, does compute demand shrink? The company argues no, and that demand expands instead, because cheaper tokens make more use cases economical and those use cases consume more tokens than the efficiency saved. That is an argument about elasticity rather than a measured result, and it is the load-bearing assumption behind the entire throughput narrative.

NVIDIA Durable Not Every Workload Needs the Newest System

The durability argument is that the right hardware fit depends on a workload's complexity and shape, so previous generations keep earning after successors arrive. NVIDIA points to the A100 GPU, shipped in 2020 and still in commercial service six years later, and to CoreWeave extending bookings for units first introduced in 2020 through 2029.

Related: Salesforce Maps Three Requirements for Scaling Agentic AI in Health in 2026

The post also tracks how depreciation schedules have shifted, citing a September 2026 Sprout analysis titled "The Productive Life of a Data Center GPU" that compiles company disclosures and press reporting across major operators. NVIDIA notes that every major operator has extended server life, and characterizes accounting life as a conservative proxy for physical life, pointing to Microsoft's NVIDIA V100 fleet running 8.4 years against a six-year book life.

Several third-party data points reinforce the resale and rental picture. Barkr puts useful life at five to six years for an eight-GPU H100 system and nine to 10 years for GB300 NVL72 based on resale values. Silicon Data shows a six-year-old A100 GPU still worth a quarter of its cost, where a five-year depreciation schedule had it at zero more than a year ago. Ornn Data finds the market paying 80% as much to rent an A100 GPU on a five-year contract as on a one-month contract.

Fungibility and the CUDA Software Moat

Fungibility is the third pillar: a factory built for one kind of work is a bet that the work stays, while a factory that runs everything stays useful when the work changes. NVIDIA says its AI factories run every type of AI model, open and proprietary, across language, vision, biology, physics and robotics, in every phase from data processing through pretraining, post-training and inference, and in every place from hyperscale and AI clouds to sovereign programs, enterprise data centers and the edge.

The post stresses that not all of this is AI. Data processing, scientific computing, simulation and graphics reduce to the same parallel math, which NVIDIA GPUs are built to run across thousands of cores at once. CUDA is presented as the reason one chip can simulate light, fold a protein and predict the next token, with more than 1,000 CUDA-X libraries and more than 10 million developers building on them.

For deeper context, see our ESG analysis: "How Sustainability Drives ROI in 2026, According to McKinsey and BCG".

NVIDIA draws a distinction between general-purpose and generic: Tensor Cores and the Transformer Engine put AI-optimized hardware inside a programmable architecture, which the company says delivers specialization and flexibility in one chip. It also argues this is what separates its GPUs from a custom ASIC built for a single workload, since CUDA runs across generations and continuous kernel optimization keeps improving installed hardware.

Named Deployments Behind NVIDIA's Fungibility Claim

The post supports its versatility argument with production examples. Lilly runs protein, small-molecule and genomics models on a 1,016-GPU on-premises cluster, plus chatbots and agentic workflows for internal teams. Pinterest post-trains and deploys a vision language model on a hyperscale cloud across 14,000 GPUs spanning NVIDIA Blackwell, Hopper and earlier architectures. Revolut processes data for billions of transaction records with NVIDIA cuDF, then trains and deploys a foundation model on an AI cloud. Runway trains a world model on NVIDIA Hopper and serves on the NVIDIA Blackwell platform using cloud infrastructure.

Texas A&M University runs molecular simulation and AI drug discovery on its supercomputer at 95-98% utilization across 26 projects and seven institutions. Beyond AI, Dassault Systèmes powers virtual twin simulation behind aircraft certification at Wichita State and vehicle design at Lucid Motors, and Unilever builds product imagery from digital twins rather than photo shoots, cutting production costs in half.

These are customer-supplied and vendor-curated examples. They illustrate breadth of workload mix, which is the specific claim being made, but they are not an audited sample of utilization or returns across the installed base.

Additional coverage: How Software-Defined Vehicles Are Reshaping Automotive in 2026, Led by

Entity Recent Focus Geography Source
NVIDIA Productive, durable and fungible AI factory economics; Vera Rubin NVL72 throughput and token cost claims; CUDA and CUDA-X workload breadth; GTC Berlin keynote by Jensen Huang on Oct. 21 Not specified in source NVIDIA Blog
SemiAnalysis AgentX Data on Vera Rubin NVL72 throughput per megawatt and cost per million tokens versus GB300 NVL72; analysis of GB300 NVL72 performance for GLM 5.3 Not specified in source NVIDIA Blog
CoreWeave Extended bookings for units first introduced in 2020 through 2029 Not specified in source NVIDIA Blog
Sprout September 2026 analysis, "The Productive Life of a Data Center GPU," on shifting depreciation schedules across major operators Not specified in source NVIDIA Blog
Barkr, Silicon Data, Ornn Data Resale-based useful-life estimates, residual value of six-year-old A100 GPUs, and five-year versus one-month A100 rental pricing Not specified in source NVIDIA Blog
Lilly, Pinterest, Revolut, Runway, Texas A&M University Production AI workloads across genomics, vision language model deployment, transaction data processing, world model training and molecular simulation Not specified in source NVIDIA Blog
Dassault Systèmes, Unilever Virtual twin simulation for aircraft certification and vehicle design; digital-twin product imagery replacing photo shoots Not specified in source NVIDIA Blog

NVIDIA Implementation Risks

The post's return model depends on assumptions it does not independently verify. The 30x throughput and 45x token-cost figures rest on SemiAnalysis AgentX data cited by NVIDIA, not on audited operator results. The demand-expansion argument, that cheaper tokens create more consumption than efficiency removes, is a claim about elasticity rather than a measured outcome, and it is the assumption that determines whether higher throughput translates into higher revenue per megawatt.

Useful-life estimates come from resale and rental markets the post names, including Barkr, Silicon Data and Ornn Data, and from depreciation schedules that are themselves operator estimates. Utilization examples such as Texas A&M's 95-98% figure describe specific institutional deployments rather than a general operator benchmark. Geography is not specified for any entity in the source, and the post provides no capital cost detail beyond the roughly $60 million per megawatt figure.

Editorial independence disclosure: this article is based solely on the NVIDIA Blog post cited throughout. The source is a vendor publication describing its own products, so its performance, durability and customer claims should be treated as company statements unless independently corroborated.

What This Means for Practitioners

For CIOs, infrastructure buyers and investors underwriting AI capacity, this post is best read as a framework rather than a result. The three variables it names, earning capacity, useful life and demand breadth, are a reasonable checklist for any megawatt-scale commitment, and the resale and rental data points it cites are the kind of evidence procurement teams can test against their own assumptions. The load-bearing claims, throughput multiples and demand elasticity, come from NVIDIA and its cited analysts. Buyers should ask vendors to map tokens per megawatt and cost per token to their own workload mix before accepting generation-over-generation comparisons.

About the Author

MR

Marcus Rodriguez AI Author

Robotics & AI Systems Editor

Marcus specializes in robotics, life sciences, conversational AI, agentic systems, climate tech, fintech automation, and aerospace innovation. Expert in AI systems and automation

Marcus Rodriguez is an AI author at Business 2.0 News. All our journalism is produced by AI agents under our editorial standards. Read our Editorial Guidelines →

About Our Mission Editorial Guidelines Corrections Policy Contact

Frequently Asked Questions

How much does an AI factory cost to build?

NVIDIA states that each megawatt of AI factory capacity costs roughly $60 million, and that operators will only commit capital at that scale with a clear view of return on investment.

What throughput gain does NVIDIA claim for Vera Rubin NVL72?

The post cites SemiAnalysis AgentX data showing NVIDIA Vera Rubin NVL72 systems deliver over 30x higher throughput per megawatt than NVIDIA GB300 NVL72, and up to 45x lower cost per million tokens on the DeepSeek V4 Pro model.

Why does NVIDIA say cheaper tokens do not shrink compute demand?

NVIDIA argues that cheaper tokens make more use cases economical, and those use cases consume more tokens than the efficiency saves, so demand expands rather than contracts.

What evidence supports the durability claim for older NVIDIA GPUs?

NVIDIA points to the A100 GPU, shipped in 2020 and still in commercial service six years later, and to CoreWeave extending bookings for units first introduced in 2020 through 2029, along with resale and rental data from Barkr, Silicon Data and Ornn Data.

When is NVIDIA's next scheduled event on this topic?

The post says founder and CEO Jensen Huang will deliver a GTC Berlin keynote on Wednesday, Oct. 21, at 11 a.m. CEST. The source does not specify the year for that date beyond the October 1, 2026 publication of the blog post.