Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/TRAINING DATA/Google's AI Infrastructure Chief…
POD
// EPISODE
TRAINING DATA

Google's AI Infrastructure Chief, Amin Vahdat, on the Physics & Economics of Frontier AI

DATE October 6, 2026SOURCE TRAINING DATAPARTICIPANTS AMIN VAHDAT, SONYA HUANG
// KEY TAKEAWAYS6 ITEMS
  1. 01Goodput, Not FLOPS, Is the Metric That Matters
  2. 02Failure Is the Normal State at Frontier Scale
  3. 03Capacity Gains Come Mostly From Models and Software, With Hardware as the Multiplier
  4. 04Specialization Is a Portfolio Bet on Workload Durability
  5. 05Co-Design Compounds Across Layers, at the Cost of Lock-In
  6. 06AI Data Centers Are Co-Designed With the Hardware, Breaking the 30-Year Fungibility Model

1. Key Themes

Goodput, Not FLOPS, Is the Metric That Matters

Vahdat argues that chip-level specs are theoretical, while what pays the bills is useful work delivered per workload under real failure conditions. At 100,000-accelerator scale, a single failure in a synchronous job can stall everything, forcing checkpoint restores and recomputation. He distinguishes throughput (work being done) from goodput (work that advances the answer). "What we really care about in the end is what's the performance delivered by workload?" 00:05:42. On accountability: "How we hold ourselves accountable is what's the good put that we deliver for the workloads that actually matter in the data center." 00:10:06. He also frames the denominator as power: "We also for us, it's good put per watt." 00:14:27. He notes the term is Google's own but spreading: "It's a Google term, but I think more and more people across the industry are starting to pick it up." 00:10:26

Failure Is the Normal State at Frontier Scale

Reliability, detection, and recovery are core competitive disciplines, not afterthoughts. Vahdat says failures are a long tail with no single root cause, and software bugs are as damaging as hardware faults. "I think that it is definitely going to be at the scale of multiple times a day. And perhaps depending on the exact configuration, multiple times an hour, something is going to fail." 00:10:39. On the absence of a silver bullet: "If there were a common failure reason, we'd have figured it out and fixed it. It is a long tail of constant discovery." 00:10:54. And on software: "You might have perfect hardware, fully reliable, but then you might have software issues that hurt you." 00:11:24

Capacity Gains Come Mostly From Models and Software, With Hardware as the Multiplier

Doubling effective serving capacity every six months is mostly not about adding FLOPS. "You have to double the capability of that hardware to generate tokens every six months. And as much or more of that is going to come from software as it is from hardware." 00:13:31. On the breakdown of intelligence per watt: "Most of the gains often come from the model side." 00:14:53. Hardware still delivers a steady multiplier: "We're living in a world right now where 2x or more year over year performance improvements is absolutely possible." 00:14:53

Specialization Is a Portfolio Bet on Workload Durability

The TPU split into 8i (inference) and 8t (training) came from a calculus of workload size and persistence. Vahdat explains: "The more you specialize to a particular workload, the less flexible it is, the faster, the more power efficient the hardware is going to be." 00:00:00. The decision hinged on inference becoming a large share of the market: "We saw inference and serving really taking off... that we thought might be 30, 40, 50, 60% of the market in its lifetime." 00:19:02. Crucially, each chip can still do the other's job, which hedges forecasting error over a six-year hardware life: "Both chips can do the other's job." 00:20:46

Co-Design Compounds Across Layers, at the Cost of Lock-In

Fully fungible, abstracted stacks leave performance on the table; co-design multiplies small gains. "There might be 10%, 20%, 2x across each of these layers. You start multiplying those optimization opportunities through and all of a sudden you're left with a big end to end opportunity." 00:24:31. The tradeoff is explicit: "Pros are you can run anywhere, anytime. You have no lock in. Cons are you're leaving significant, in all likelihood, significant amounts of performance on the table." 00:24:59. Google's advantage is a tight loop with DeepMind where chip architecture can be altered "in flight" before tape-out 00:28:59.

AI Data Centers Are Co-Designed With the Hardware, Breaking the 30-Year Fungibility Model

Traditional data centers were planned for 25-30 year horizons with 6-year hardware cycles. AI buildings are different. "An AI data center likely is going to be much more purpose built, co-designed with the hardware, even to the point of cooling and power distribution." 00:04:13. The driver is rack power density: storage racks run 10-40 kW versus "hundreds of kilowatts today" for AI racks, potentially megawatts 00:03:20.

Long-Horizon Agents Reshape the Whole Data Center, Not Just the Accelerator

Agents remove the human rate-limiter and shift load onto CPUs, memory, storage, and network. "What went from, you know, seconds, maybe tens of seconds in terms of interaction time is not going into perhaps milliseconds." 00:35:06. Vahdat adds: "All that reasoning and all that parsing is going to probably happen on a CPU." 00:35:32. Result: "The demand for accelerated computers going up, but the demand for CPU and networking and storage... is also going through the roof." 00:36:08. This creates a hard design question about co-locating heterogeneous racks versus linking buildings with higher latency 00:36:28.

Power Is the Fundamental Long-Term Constraint

All constraints are hard and shift, but power is the deepest. "If I had to answer fundamentally, I would say that power is the single most fundamental constraint that we face. Like everything else seems like we know how to solve them." 00:42:47. Google's preference is to partner with utilities, covering the transmission upgrade costs, and optionally bring local generation that can feed back to the grid: "By leveraging that statistical multiplexing over a much larger base, actually everybody wins." 00:47:02

Optical Circuit Switching as a Reliability and Flexibility Weapon

Google has used optical circuit switching (MEMS mirrors) for ~15 years. In TPU clusters it lets the system swap a failed rack for a spare in milliseconds with no human rewiring. "When a rack fails, we redirect the light to that new rack. Huh. And that can be done in milliseconds." 00:41:24. This ties directly to goodput: fast failover protects the metric that matters.

Open Standards as Strategy

Despite deep vertical integration, Google supports interoperability (e.g., Torch TPU alongside JAX). Vahdat invokes history: "Why did the internet protocol win?... it was this narrow waist to the hourglass where anything could run on top in terms of software and anything could run underneath it in terms of hardware." 00:54:53. And: "It's gotta be one of choice." 00:56:14

2. Contrarian Perspectives

Specialization "Never Wins" Was Wrong, but Only Conditionally

The 2013 consensus held that custom accelerators lose to Moore's Law and general-purpose programming. Vahdat: "All the smartest, wisest people would say, you don't build a custom built accelerator for a single workload." 00:15:45. And Sonya's framing: "the bitter lesson of chips" met his reply, "specialization never wins" 00:16:17. The nuance is that specialization wins only when you can predict workload durability and size, and he notes "there were quite a few people even within the company... who are not sure that it would work out." 00:16:17

Eight-Year-Old Chips Are Still Fully Utilized, Undercutting the Short-Depreciation Narrative

In the live debate on the useful life of AI chips, Vahdat says: "I said that our seven and eight-year-old TPUs are still at 100% utilization." 00:52:16. Yet he also says replacement still makes sense after the ~six-year depreciation due to power efficiency, so the answer isn't "chips die quickly" or "chips last forever" but a power-per-token replacement calculus 00:52:30.

Training Clusters Can't Simply Be Recycled Into Inference

The tidy framework is that you train on the newest cluster, then cascade it to inference. Vahdat pushes back: "The demand for inference is such that we can't just rely on whatever older training clusters that are no longer being fully used just for training as the basis for inference." 00:49:46. Inference needs geographic distribution and a mix of compute, networking, and storage, so it requires purpose-built sites on other continents 00:50:14.

Inference Datacenters Are Not Cheaper Per Megawatt

The intuition is that serving is the commodity tier. Vahdat: "Not necessarily cheaper, actually. Because for serving, there is more need to co-locate compute networking and storage. And so then we get to the inability to specialize." 00:50:58. Training gets uniform density; serving is heterogeneous and distributed.

Cooling in Space Is Harder, Not Easier

Orbital compute is pitched on abundant energy, but Vahdat flags the naive assumption: "Naively, you might think cooling in space is easier. It's actually harder." 01:00:01. He also notes the upside is quantifiable: roughly 1.4x solar capacity and 90-100% sun coverage versus 28-35% on land 00:59:14, and that "there are no showstoppers here, no fundamental showstoppers." 01:00:33

3. Companies Identified

Google / Google Cloud

Vahdat's employer; hyperscaler running the largest AI infrastructure build-out. Mentioned for spending "more than $200 billion on CapEx this year" (Sonya Huang) 00:00:59 and for offering customers choice of TPUs, GPUs, and other accelerators. Vahdat: "What we like to do at Google is give our customers choice." 00:23:18

Google DeepMind

Google's frontier AI lab, a co-design partner on hardware. Vahdat: "It's one of the most fun and frankly gratifying parts of being at Google is the opportunity to work really shoulder to shoulder with the DeepMind team in terms of co-design of our hardware and models." 00:26:23. He adds that Gemini is used to design hardware for future Geminis: "We're using Gemini to design hardware for future Geminis as well." 00:30:26

NVIDIA

GPU vendor and partner. Vahdat praises it as more than a chip company: "NVIDIA is an incredible whole systems company. I mean, they're obviously a semiconductor company, but they're not just a semiconductor company. They give you a really, really strong reference stack." 00:12:13. Google sells and uses GPUs and delivered a Vera Rubin cluster to Ineffable Intelligence.

Ineffable Intelligence

AI company (Sonya Huang's portfolio company) that received a big Vera Rubin cluster from Google. Vahdat: "Fantastic team. That got a lot of attention." 00:04:49. The fiber-cabling photos were among Google's most popular posts ever 00:05:14.

OpenAI and Anthropic

Frontier labs referenced for their differing compute strategies. Sonya framed OpenAI as primarily homogeneous and Anthropic as more heterogeneous; Vahdat declined to speculate on their architectures: "I can't speculate... I would imagine that there could be many reasons for why they wind up with different architectures." 00:25:55

JAX / PyTorch (Torch TPU)

Model frameworks. JAX is Google's own; PyTorch is widely used by customers. Vahdat uses these as the example of open interoperability: "If you like JAX, you can use it. But if you like PyTorch, we have Torch TPU where your unmodified models, et cetera, can run." 00:54:53

Startups building model-specific chips (unnamed)

Vahdat acknowledges a further level of specialization beyond transformers: "You could go further and specialize to not just the transformer, but your model... I think there are a number of companies out there that are thinking about that. I think it's a very, very interesting direction as well." 00:22:17

ISPs and Utilities (PG&E referenced)

Partners for edge sites and power respectively. Vahdat: "We actually partner with ISPs. That might be a rack." 00:48:51. On utilities (PG&E raised by Sonya): "Our preferred model at Google is always to be utility connected, to be good connected." 00:43:58

4. People Identified

Amin Vahdat

Head of Google's AI infrastructure (named at the end of last year), leading one of the most capital-intensive build-outs in history. Mentioned as the guest and operator. Sonya Huang: "You are the man spearheading the efforts of one of the most capital intensive build outs in human history." 00:00:59

Demis Hassabis

Leader of Google DeepMind (referenced by first name only, "Demis"). Mentioned as a frequent counterpart. Vahdat: "I talk to whether it's Cori or Demis multiple times a week, etc." 00:29:53

Koray Kavukcuoglu

Senior Google/DeepMind AI leader (transcribed as "Cori"). Mentioned alongside Demis as a regular collaborator. Vahdat: "I talk to whether it's Cori or Demis multiple times a week." 00:29:53

Sonya Huang

Investor and host; her portfolio includes Ineffable Intelligence. Mentioned as the interviewer who raised the orbital compute, goodput, and depreciation questions.

5. Operating Insights

Build Hardware "Intercept" Windows Into Your Org Structure

Because chips run on multi-year pipelines, Google maintains parallel generations at five or six stages, and the model team stays in the loop at each. Vahdat: "We have the chips that are in production... in implementation that are about to tape out... in design phase... and then the chips that are in concept phase. So it's really this five, six stage, many, many year pipeline." 00:28:06. The tactic is to intercept at the right stage: "We can make changes to the chip, literally chip architecture in flight, which would be somewhere between hard and impossible to do if we were working across company boundaries." 00:28:59

Set a Materiality Bar for Cross-Team Disruption

Researchers self-filter requests to hardware: they don't ask for 1% changes because they know stopping a tape-out is costly. Vahdat: "If they have a small tweak that's going to deliver like 1% or 0.5%... they're probably not going to come to us because they know... it's not like software where you can just do a change list." 00:31:24. A big opportunity, though, triggers days of joint work and possibly delaying a tape-out by "a week or two weeks" 00:28:06. Pairing shared context with a self-enforced materiality threshold keeps the interface efficient.

Negotiate to the 90% Solution Jointly

When a model idea can't be fully supported in hardware, the teams trade off both sides: "Maybe you can change your model architecture in this other direction that gives you 98%, that gives us 90% of what we're looking for." 00:27:43. Neither side demands the ideal; they co-optimize.

Use AI Internally Beyond Software: Hardware Design and Site Planning

Vahdat reports hardware engineers use AI as much as software engineers: "Our hardware engineers are using AI as much as the software engineers are. And productivity has gone up. Time from design kickoff to tape out is shrinking." 00:56:57. In data center site planning, AI aggregates information for humans rather than replacing them: "Putting the right information in front of the humans making the decisions." 00:58:12

Pay for the Externalities When Securing Power

To maintain utility goodwill and avoid raising residential rates, Google covers the transmission and substation costs it triggers, and plans many years ahead. Vahdat: "We ensure that... the transmission lines that have to be built, upgraded, additional utility base stations, et cetera, that we pay for those as well." 00:44:38

6. Overlooked Insights

The TPU's Core Architecture Has Barely Changed Since V1

In passing, Vahdat says the architecture has been remarkably stable at a medium level of detail across generations of model algorithms. This is the real reason five-year hardware forecasting is tractable: "The TPU architecture at a medium level of detail... hasn't really changed since TPU V1." 00:32:54. A small set of primitives (large matrix multiply units, a sparse core for vector, scatter-gather operations, remote load/store over ICI) has stayed durable, which helps explain why Google can commit to long hardware roadmaps without predicting algorithms 00:33:39. The implication for investors is that durable primitives, not speculative workload prediction, are the moat in AI silicon.

Sync Workloads Make Network Reliability a Direct Revenue Lever

Buried in the goodput discussion is the observation that essentially all major workloads (training, serving, agentic) are synchronous, so the blast radius of a single fault is the entire job 00:07:04. Combined with the remark that telemetry and fast failure detection is "like finding a needle in the haystack continuously" 00:10:06, and optical circuit switching enabling millisecond rack replacement 00:41:24, it implies that fault detection, observability, and rapid failover tooling are high-leverage, under-discussed investment areas. At 100,000-accelerator scale, every percentage point of goodput recovered is effectively free capacity.