Why smarter AI models could drive up compute prices 10x
- 01The Revenue-Compute Gap Is Structurally Unsustainable Without a Pressure Release Valve
- 02Anthropic's Revenue Trajectory Points to Trillion-Dollar Scale by End of 2026
- 03Inference Margins Have Already Roughly Doubled in Under a Year
- 04Compute Price Inflation Is Already Underway, Especially for Frontier-Grade Supply
- 05The True Value of an H100 Should Be 15x Higher If AI Reaches Human-Level Software Engineering
- 06The Labs Have a Strategic Incentive NOT to Shift Compute to Inference
Dwarkesh Patel | Dwarkesh Podcast
1. Key Themes
The Revenue-Compute Gap Is Structurally Unsustainable Without a Pressure Release Valve
Dwarkesh identifies a fundamental divergence: AI lab revenue is growing 10x per year while compute capacity only grows 3x. This gap must be resolved by some combination of margin expansion, compute price increases, or shifting compute from training to inference.
"For a lab to keep 10x-ing revenue year over year while compute only 3x, one of the following three things needs to happen... One, lab margins have to increase. Two, the price of compute has to increase. Or three, the percentage of compute that lab spent on inference rather than training has to increase." 00:00:27
Anthropic's Revenue Trajectory Points to Trillion-Dollar Scale by End of 2026
Three consecutive years of 10x revenue growth creates a compounding projection that is staggering: $9B at end of 2024, potentially $100B–$150B end of 2025, and $1T by end of 2026. Dwarkesh is candid that this depends entirely on whether AI capabilities actually deliver that much value.
"Anthropic's revenue has 10x year over year, and it's likely to do so again this year. So they ended last year with $9 billion in revenue. I think they'll probably end this year with somewhere between $100 billion to $150 billion in revenue. Now, for this trend to continue, Anthropic would need to make $1 trillion in revenue by the end of next year." 00:00:00
Inference Margins Have Already Roughly Doubled in Under a Year
This is a concrete, specific data point on how quickly the economics are shifting — Anthropic's inference margins went from 40% to upward of 80% within roughly 12 months. This is an enormous operational transformation happening quietly.
"With regards to the margins, Anthropic's inference margins reportedly went from 40% in the middle of last year to upwards of 80% now available." 00:00:55
Compute Price Inflation Is Already Underway, Especially for Frontier-Grade Supply
Google paying 2x spot price for reserved GPU clusters, combined with spot prices themselves being 40%+ above their February 2025 trough, signals that the compute price escalation thesis is not theoretical — it is already in motion.
"Google, for example, is paying $900 million a month for 110,000 GPUs that are a blend of GB200s and GB300s. The price that Google is paying here is 2x the spot price per hour for those GPUs. And that spot price itself is more than 40% higher than it would have been in February." 00:03:32
The True Value of an H100 Should Be 15x Higher If AI Reaches Human-Level Software Engineering
This is the most striking valuation reframe in the episode. If a single H100-equivalent can run a true human-level software engineer, the implied rental value of that chip — at today's software engineer salaries — would be over $250K per year, versus current spot pricing.
"If a true human-level software engineer could run on an H100 equivalent, then at today's prices for software engineers, that H100 should rent for over $250K a year. That's over 15x the current spot price for an H100. And this is not even accounting for the fact that your AI can work nights and weekends." 00:04:00
The Labs Have a Strategic Incentive NOT to Shift Compute to Inference
Despite the revenue pressure, labs deliberately resist allocating the majority of compute to inference because doing so would signal that AI progress has stalled and recast them as commodity cloud providers rather than AGI builders — a strategically and narratively inferior position.
"The way the labs see the world, the whole point of inference revenue is to help convince investors to give you more money in order to train the next bigger, better model. And if you're spending most of your compute on inference, you're basically declaring that AI progress has stalled and you're just now in the business of being a cloud provider." 00:01:21
The Alkin-Allen Effect Means Model Efficiency Will Command a Premium at Scale
As compute becomes more expensive, the economic penalty for using an inefficient model compounds. Labs that can train models requiring fewer tokens to achieve the same result will be able to charge dramatically higher margins — essentially creating synthetic compute through efficiency.
"If it costs $20 an hour to rent an H100, then it would be extremely stupid to use a weaker, less efficient model because it's going to burn more tokens on your expensive compute to get the exact same result. So labs will be able to charge a much larger premium if they can train a model that better economizes this scarce input." 00:05:49
Consumer AI Applications Will Get Priced Out as Compute Becomes More Valuable
As the marginal value of compute rises, frontier lab use cases — particularly AI doing AI research — will outbid consumer applications like AI content generation. This is a coming displacement of the current democratized-access moment.
"Google or Anthropic or OpenAI will be willing to pay more for the tokens to automate AI research than you or I will be willing to pay to make more AI slop talk." 00:06:16
The 3x Annual Compute Scaling Rate Is Already Near Its Ceiling
Dwarkesh breaks down the 3x into its three distinct components — Moore's Law (1.4x), new fab construction (1.2x), and AI absorbing wafer share from smartphones/PCs (1.8x) — and argues that each component faces hard near-term ceilings, making even sustaining 3x a challenge, let alone accelerating it.
"1.8x comes from the fact that AI is absorbing a lot of wafer allocation that was previously going to smartphones and PCs. This is probably going to hit a wall by the end of next year when at the leading edge N3 nodes at TSMC, AI will have gone from 60% to 86%. At some point, you have just absorbed all leading edge wafer capacity for AI." 00:08:11
Strong Economies of Scale in the Model Business Favor Dangerous Power Concentration
Unlike human labor, a trained model's skills are shared across all users from a single training run. This creates a structural economy-of-scale advantage unlike anything in prior technology, and Dwarkesh openly worries about what it means for power concentration.
"When you train a model, you just had to spend this one-time cost to learn all these different skills that then get to be shared across all your users. This is very unlike human labor, where each instance has to be retrained from scratch. I wish we didn't live in a world with such strong economies of scale for intelligence, because I'm worried about power concentration." 00:10:31
2. Contrarian Perspectives
>90% Inference Margins for Intelligence Are Plausible — and May Not Get Competed Away
The conventional view is that margins this high invite competition and get eroded. Dwarkesh pushes on whether the leading model might be so decisively better than alternatives that competition cannot close the gap, making extreme margins durable. He presents this as uncomfortable but structurally possible.
"It's just really wild for me to consider that the margins for something like intelligence will be greater than 90% and they don't get competed away at that level." 00:02:39
The Lump-of-Labor Fallacy Applies to AI — More AI Workers May Not Crash Wages
The intuitive reaction to 10 million new AI software engineers appearing would be: wages collapse. Dwarkesh invokes standard labor economics to argue this could be wrong — innovation and specialization historically maintain the marginal value of labor even after supply shocks, so compute value may stay astonishingly high.
"If we apply this argument to people instead of AIs, then this would be the classic lump of labor fallacy. Economists generally believe that high-skill immigration does not decrease wages in the long run because of how innovation and specialization increase the value of labor." 00:04:28
The Simon-Ehrlich Bet Analogy Understates AI Compute Scarcity
The famous bet is used to argue against scarcity predictions. Dwarkesh pushes back: commodity supply is elastic and substitutable; compute supply is not. The bottlenecks — ASML EUV machines, fab construction timelines, wafer allocation ceilings — have no easy substitutes and are not responsive to price signals at AI timescales.
"I think the supply of compute is much less elastic and much less capable of absorbing large demand shocks and much less capable of being accommodated by using different substitutes than the extraction of different metals." 00:07:14
Labs Training on Sparse Compute Will Be Competitively Disadvantaged Against Labs That Can Monetize Expensive Compute
As compute becomes scarcer and more expensive, labs that train more efficient models gain a compounding advantage — they get better results per dollar AND can charge higher premiums. This inverts the typical assumption that cheaper models win on accessibility.
"If you have a model that can get the same result by using less compute, then you've in some sense created more compute, and the value of compute is going to increase." 00:05:49
3. Companies Identified
Anthropic
AI safety-focused lab, creator of the Claude model family. Cited as the primary case study for the 10x annual revenue growth trend and the dramatic improvement in inference margins (40% → 80%+). The entire analysis is anchored on Anthropic's publicly reported financials.
"Anthropic's revenue has 10x year over year, and it's likely to do so again this year. So they ended last year with $9 billion in revenue." 00:00:00
Cited as a concrete data point on premium compute pricing: paying $900M/month for 110,000 reserved GPUs (GB200/GB300 blend) at 2x the already-elevated spot price, illustrating frontier lab willingness to pay for guaranteed, secure, large-scale compute.
"Google, for example, is paying $900 million a month for 110,000 GPUs that are a blend of GB200s and GB300s. The price that Google is paying here is 2x the spot price per hour for those GPUs." 00:03:32
SpaceX
Mentioned as the compute provider from whom both Google and Anthropic are renting GPU clusters, establishing SpaceX as a significant player in the frontier compute supply chain.
"A relevant case study here is to look at the compute that Google and Anthropic are renting from SpaceX." 00:03:32
OpenAI
Cited via Epoch AI data showing that in 2024 OpenAI was spending only 25% of its compute on inference — a benchmark for how the training/inference compute allocation ratio is shifting across the industry.
"In 2024, according to Epoch, OpenAI was spending just a quarter of its compute on inference. And that number is likely closer to 50%, if not higher now." 00:01:21
ASML
Identified as the single hardest bottleneck to compute scaling: the supplier of EUV lithography machines required to build new leading-edge fabs. The 1.2x annual contribution from new fab construction is explicitly constrained by ASML's machine production rate through 2030 and potentially beyond.
"This process is ultimately going to be bottlenecked up to 2030 and potentially even beyond by just building new ASML EUV machines." 00:08:11
TSMC
Cited as the leading-edge fab operator where the wafer-absorption dynamic is most visible: AI has grown from 60% to an expected 86% of N3 node wafer allocation, fast approaching a hard ceiling.
"At the leading edge N3 nodes at TSMC, AI will have gone from 60% to 86%." 00:08:11
Mercury
Dwarkesh's banking platform, specifically its AI feature "Command," which automates transaction categorization by reasoning over vendor, purchaser identity, and memos — syncing output to QuickBooks.
"I have Command, which is Mercury's built-in AI. Take a stab at all of them at once. Command proposes a category for each transaction and provides its rationale." 00:09:08
4. People Identified
Dylan (guest, prior episode)
Referenced as having provided detailed analysis of the ASML EUV bottleneck and fab construction constraints on compute scaling on a prior Dwarkesh podcast episode. His last name is not given in the transcript.
"Dylan, when he was on the podcast a few months ago, talked about this in great detail." 00:08:11
Paul Ehrlich
Biologist famous for population doomsday predictions; lost the Simon-Ehrlich bet on commodity prices. Cited as the cautionary analog for scarcity predictions — but Dwarkesh argues the analogy ultimately breaks down for compute.
"Paul Ehrlich was this famous doomer about population growth, and he made this bet that a basket of commodities would increase in price rather than decrease in the decade preceding 1990." 00:06:45
5. Operating Insights
Model Efficiency Is Now a Direct Revenue Multiplier, Not Just a Cost-Saver
For any company building on top of AI APIs, the Alkin-Allen dynamic means that as compute prices rise, the efficiency of the underlying model directly affects unit economics. Operators should evaluate models not just on output quality but on tokens-per-result — and should expect to pay substantially higher premiums for demonstrably efficient frontier models as compute tightens.
"If it costs $20 an hour to rent an H100, then it would be extremely stupid to use a weaker, less efficient model because it's going to burn more tokens on your expensive compute to get the exact same result." 00:05:49
Locking in Long-Term Compute Contracts Now Has Asymmetric Upside
The evidence that reserved/contracted compute already trades at 2x spot — and that spot itself is 40%+ above its February trough — means that operators with meaningful inference workloads should treat long-term compute contracts as a strategic hedge, not just a procurement decision. Waiting for spot markets to normalize is probably the wrong posture.
"Google, for example, is paying $900 million a month for 110,000 GPUs... The price that Google is paying here is 2x the spot price per hour for those GPUs. And that spot price itself is more than 40% higher than it would have been in February." 00:03:32
Consumer-Facing AI Products Built on Cheap Inference Are in a Competitively Fragile Position
If the marginal value of a token is going to be bid up by frontier lab workloads, low-margin consumer AI applications will face margin compression from both sides — rising input costs and difficulty passing those costs to price-sensitive users. Founders in this space should be building toward either high-value enterprise use cases or achieving deep enough efficiency gains to survive repricing.
"A lot of current popular applications of AI will probably get priced out... Google or Anthropic or OpenAI will be willing to pay more for the tokens to automate AI research than you or I will be willing to pay to make more AI slop talk." 00:06:16
6. Overlooked Insights
The Training/Inference Reallocation Is a One-Time, Non-Repeatable Lever That Is Almost Exhausted
Dwarkesh notes in passing that AI's absorption of wafer share from smartphones and PCs contributed 1.8x to annual compute growth — the single largest component of the 3x. But this is structurally a one-time reallocation: once AI reaches ~86% of N3 leading-edge wafer allocation at TSMC by end of 2026, this lever disappears entirely. This means the effective compute growth rate will drop sharply and permanently, not gradually, creating a sudden step-down in available scaling. This is arguably more significant than any individual GPU pricing data point, yet it receives only a single sentence of attention.
"1.8x comes from the fact that AI is absorbing a lot of wafer allocation that was previously going to smartphones and PCs. This is probably going to hit a wall by the end of next year when at the leading edge N3 nodes at TSMC, AI will have gone from 60% to 86%." 00:08:11
The Post-Singularity Compute Repricing: Robots Converting Sand Into Chips Is the Actual Endgame
Briefly and almost dismissively mentioned, Dwarkesh points to a future where AI-driven robots can manufacture chips from raw materials (silica sand, copper), at which point compute's price collapses to raw input costs. This is the actual resolution of every constraint described in the episode — and it implies a specific investment thesis: the companies building physical AI and robotic manufacturing automation are not just interesting; they are the entities that will eventually break every bottleneck discussed here. This was stated in a single transitional sentence and not discussed further.
"At some point, we'll just have robots that can convert shores of silica sand and mines of copper into new computer chips. And then the price of compute is basically the raw inputs and the tools required to do this processing." 00:10:05