Chips, Memory, and Power | Pat Gelsinger
- 01AI Makes Chip Design Easy, So the Bottleneck Moves to Manufacturing and Packaging
- 02The Unit of Compute Is No Longer the Chip
- 03The Hundred AI Chip Startups Will Consolidate
- 04Memory Innovation Is Finally Coming After 30 Years of Stagnation
- 05Stacking Has a Yield Ceiling: Mid-Rise, Not Skyscraper
- 06Optics Everywhere Except the Core Compute-Memory Complex
1. Key Themes
AI Makes Chip Design Easy, So the Bottleneck Moves to Manufacturing and Packaging
Gelsinger compares today's AI-driven design shift to the 486 era, when engineers invented EDA tooling because it did not exist. His argument is that fast design only exposes slower downstream steps. He says "I can design the thing in three months, but I can't actually get it into real silicon at scale for nine months" 00:10:19, and breaks the nine months into roughly three months of fab time, three months of advanced packaging, and then rack-scale integration. His reason this matters: "if it takes me a year and a half until I actually get scale and software in it, okay, my understanding of the AI workloads is no longer applicable to the chip that I design" 00:11:29. He wants new lithography and design flows that avoid "$50 million of mass cost until I can get things into prototyping."
The Unit of Compute Is No Longer the Chip
Gelsinger states plainly that "nothing's a chip anymore, it's a rack" 00:00:00. The packaging, 3D stacking, power delivery, optics and rack integration are now as important as the silicon. That is why he sees AI compressing design time while leaving system-level timelines largely untouched.
The Hundred AI Chip Startups Will Consolidate
Gelsinger expects the field to narrow for three reasons. First, workload heterogeneity (pre-fill, mid-fill, decode, reasoning that wants CPU-like behavior) is unstable: "I really don't like, I'll say, at-scale specialization, right? Because I think the workloads are going to continue to moderate, you know, and migrate so significantly" 00:15:55. Second, scale requires capital and workloads, which not all can win: "you have to get scale on these things. And scale requires, you know, you have to win, you have to get capital, you have to get workloads onto it" 00:17:46. Third, the biggest players will pick winners: "the winners will pick some winners... an OpenAI, an NVIDIA, an Anthropic will say, I like this one because it isn't just hardware. It's how do the hardware and software co-evolve." He also notes the incumbent absorption pattern: "we saw this with what NVIDIA did with Groq... two years from now, you're not going to tell what Groq is inside of NVIDIA" 00:21:21.
Memory Innovation Is Finally Coming After 30 Years of Stagnation
Gelsinger frames memory as the most under-appreciated physics problem: "Exactly how many major new memories have we had over the last 30 years? Zero" 00:00:27. What changed is the economics: "AI is a memory workload" and "the combination of capital and technology needs that will justify memory innovation" 00:24:33. He calls HBM "a hideous memory," with poor bit density and limits on shoreline bandwidth, power and thermals, and says "the memory industry has now increased by market cap by $2.5 trillion over the last four years" 00:27:08. He has "just funded a new memory company as well. That's still in stealth," and names ferroelectrics as one materials direction.
Stacking Has a Yield Ceiling: Mid-Rise, Not Skyscraper
Gelsinger is skeptical of 16- or 32-high stacks: "the yield of any individual layer in the stacking process has to keep... go up exponentially" 00:27:50. He thinks two-to-four-high memory stacks, with chiplets for yield and thermals, will be the sweet spot: "I'm not a crazy 16, 32 stack kind of guy. I think there's going to be nice three, four, five stacks" 00:28:39. He adds power rails (the Z dimension) and RDL layers that bring a canonical package to "about eight layer stick." On thermals: "We engineers are becoming plumbers" 00:31:54, with diamonds, liquid cooling and better power regulation among the options.
Optics Everywhere Except the Core Compute-Memory Complex
Gelsinger is "a huge fan of moving to optics for all IO function" but opposes optical links between compute and memory because of energy cost: "the femtojoules per compute is now actually pretty amazingly good... The femtojoules per bit of communication is comparatively a thousand times worse" 00:33:10. He sees copper becoming uneconomic at scale-up distances: "to make my copper work at five meters is more expensive than making optical work at 100 meters" 00:35:47. He puts the inflection at 2028-29 via NPO/CPO because "I can't scale up to large radix compute clusters for my scale-up environment without making that move" 00:36:16. He also expects optical circuit switching: AI flows are "large and predictable," so "that looks more like a switch than like a packet" 00:38:39, which dissolves the scale-up/scale-out boundary.
Energy Is the Binding Constraint on AI Capex
Gelsinger links economic capacity directly to power: "in an AI digital age, energy capacity equals economic capacity" 00:00:00. He argues the U.S. was effectively flat on capacity for 15 years and has only recently reached roughly 4% annual growth. He predicts data center defaults: "I think you're going to see more and more defaults happening on many of those data center projects because the energy won't be there" 00:45:21, and cites Oracle as the first indicator. Supply-chain constraints include gas turbine lead times he puts at eight years and renewables with "deep dependencies in China."
Power Delivery Is Being Rebuilt from the Grid to the Die
Gelsinger lists a stack of innovation opportunities: 800-volt DC data centers ("the revenge of Edison is upon us" 00:46:36), vertical GaN converting 800V directly to 48V, 12V or 5V in one step, solid state transformers and switches, and cooling systems (turbines, refrigerants). He notes that on-chip power modeling is crude: "I'm essentially guard banding almost my full power envelope... I'm losing 40% of my power in guard banding" 00:12:58, and that CMOS efficiency per teraflop has flatlined "for the last five generations of NVIDIA chips." His response is fundamental physics change, including superconducting chips.
Training and Inference Will Converge via Continuous Learning
Gelsinger argues that heterogeneity is less likely because training and inference will blur: "I do think the separation of training and inferencing is going to become less so going forward, not more so... continuous learning becomes the nature of the algorithm where, yeah, I do want my inference environment to be updating my model weights" 00:40:35. He also predicts a return of HPC-style 64-bit precision workloads as AI models drive chemical and biological simulations.
Virtualization Must Be Rebuilt for Agents
Gelsinger sees every VMware-era function reappearing for agents: "Who's going to manage the agents? Who's going to create the security profiles around all of the agents? Who's going to manage the performance of the agents? ... every fundamental element of virtualization and management needs to get recreated in this next computing hierarchy" 00:50:10. He imagines "V-motion for agents," with humans setting policies and "constitutions" through dashboards, while the hard abstraction work serves agents rather than hardware and humans.
2. Contrarian Perspectives
The Explosion of AI Chip Startups Is a Temporary Phase
Most investors and founders treat the hundred-plus inference accelerators as a durable new heterogeneous era. Gelsinger, who has funded his share, says it is not: "I see it as a more temporary thing and it will converge is what I expect" 00:14:59. His backup is the specifics of the market: workloads migrate faster than a 1.5-year chip cycle can track ("Graphcore wasn't a bad design, but the world moved on" 00:11:45), scale requires billions in capex ("could I borrow 50 billion?" 00:21:21), and customers like OpenAI, NVIDIA and Anthropic will anoint a few partners. He grants that heterogeneity will persist but be "abstracted underneath" and hidden inside large vendors.
Processing-in-Memory and Optical Memory Disaggregation Are Both Wrong Bets
Two popular ideas for breaking the memory wall get rejected. On PIM: "most PIM ideas have been around for 25, 30 years. And I think 25, 30 years from now, they'll still be hanging around" 00:34:31, because they constrain workloads. On pushing memory off-package over optics: "I burn a lot of power getting there... I think I somewhat defy the underlying physics of the solution" 00:33:10. His preferred answer is boring and physical: denser, higher-bandwidth memory placed close to compute, via new memory physics.
Taller Chip Stacks Are Not the Future
The prevailing narrative says 3D stacking continues to scale vertically. Gelsinger argues yield math caps it: individual layer yield must improve exponentially as the stack grows, and a "cracked die" cannot be rescued by redundancy. His call is that two-to-four-high memory stacks are the sweet spot, even though a full package ends up at about eight layers once power and RDL are included.
Energy, Not Chips or Capital, Will Cap the AI Buildout
Gelsinger asserts that the thing most likely to deflate AI euphoria is not model quality or GPU supply but power. "Why build a new data center and buy the million GPUs if I can't power them" 00:00:00 and "you're going to see more and more defaults happening" are specific, near-term claims about project-finance failures. He points to the Oracle situation as "the first of what you're going to see many" 00:45:37 and to nuclear: "When was the last nuclear reactor that came online in the U.S.? ... It was 20 years ago" 00:46:07.
The VM Abstraction Is Not Dead; It's the Template for the Agent Era
In a world where many assume containers and serverless replaced VMs, Gelsinger says "the virtual machine abstraction, whether it's done from the infrastructure level or from the application level, I think it's still a foundational abstraction model that deserves to always have a place in the compute hierarchy" 00:49:14, and that the VMware playbook (security, performance, migration, management) must be rebuilt for agents.
3. Companies Identified
D-Matrix
AI inference chip startup bringing memory and compute together in new structures. Gelsinger says Playground has this company in the portfolio and cites it as an example of attacking HBM's shoreline bandwidth limitation: "obviously, you know, we have a company, D-Matrix, that's doing this, but I think others will be, you know, whether it's Cerebras and their techniques" 00:25:14.
Cerebras
Wafer-scale AI chip company. Mentioned as another example of bringing memory and compute together in fundamentally new structures to get around HBM's limits. Quote as above 00:25:14.
NVIDIA
Dominant AI GPU vendor. Gelsinger cites its move to absorb Groq as the template for how big players "own some of that heterogeneity," and says NVIDIA has "indicated" that co-packaged optics is the conversion point for scale-up: "you're not going to tell what Groq is inside of NVIDIA, right? It's just going to be like another special" 00:21:21. He also references the NVL72 as "an engineering marvel and it's a manufacturing nightmare" 00:36:47.
Groq
AI inference chip company that NVIDIA absorbed. Mentioned as the example of incumbents consolidating specialization: "what NVIDIA did with Groq. Okay, you know, we're going to sort of own some of that heterogeneity" 00:21:21.
Graphcore
AI chip startup. Gelsinger uses it as a cautionary tale about workload drift: "Graphcore wasn't a bad design, but the world moved on, right?" 00:11:45.
Snowcat (superconducting chip company)
A Playground portfolio company building superconducting compute. Gelsinger says it "affords the opportunity to fundamentally change the physics. You know, a thousand times better power performance" 00:48:26.
Alva (nuclear)
Portfolio company in nuclear energy, cited as part of the baseload power answer: "we talked about nuclear operating as one of those with Alva" 00:44:24.
Vertical GaN power company (unnamed)
A Playground portfolio company building vertical gallium nitride power conversion. Gelsinger says "vertical GaN, I think will be a killer technology, to be able to do 800 to 48 or 800 to 12 or even 800 to 5 in one conversion step" 00:47:00.
Unnamed stealth memory company
Gelsinger says "I've just funded a new memory company as well. That's still in stealth, but I'm quite excited about some of the innovations that will occur in new memories" 00:25:36.
Memory Vendors (SK Hynix, Samsung, Micron, implied)
Referred to collectively as "all three big memory vendors" now in the top 20 most valuable companies. Gelsinger says the memory industry "increased by market cap by $2.5 trillion over the last four years" 00:27:08, and that AI as a memory workload justifies renewed R&D.
OpenAI
Mentioned for its view on speculative decode workloads and as a likely chip-partner kingmaker: "Well, maybe speculate and verify if we listen to OpenAI here" 00:15:27; "an OpenAI, an NVIDIA, an Anthropic will say, I like this one" 00:18:43.
Anthropic
Cited as one of the large buyers that will pick hardware winners, and for Dario Amodei's constitution concept, which Gelsinger borrows for agent policy 00:51:39.
Oracle
Cited as the first indicator of data center projects hitting energy-constrained default risk: "the Oracle, you know, was just, I think that's the first of what you're going to see many" 00:45:37.
Cadence/Synopsys-style EDA industry (via Berkeley/Intel origins)
Gelsinger notes that Intel's 486 work with Berkeley researchers helped usher in "many of the foundations of what became the modern EDA industry" 00:07:30.
VMware
Gelsinger's former company, whose abstraction model he argues should be recreated for agent management: "every fundamental element of virtualization and management needs to get recreated in this next computing hierarchy" 00:50:10.
Playground Global
Gelsinger's firm and the source of several named portfolio companies (D-Matrix, Snowcat, Alva, vertical GaN company, stealth memory company).
Intel
Gelsinger's former employer and the setting for his early chip design work on the 286, 386 and 486, and his contention that "we ushered in many of the foundations of what became the modern EDA industry" 00:07:30.
4. People Identified
Pat Gelsinger
General partner at Playground Global, former CEO of Intel and VMware, former Intel CTO. He started at Intel at 18 and was engineer number four on the 386 and architect and design manager for the 486. His thesis throughout is that bottlenecks are moving from design to manufacturing, memory, power and optics. "Starting as a technician, moved into the design team at the end of the 286, engineer number four on the 386, architect and design manager for the 486" 00:03:50.
Raghu Raghuram
A16Z partner, former VMware CEO, host and Gelsinger's former colleague. He drives the discussion on memory innovation, stacking, optical memory and power.
Guido Appenzeller
A16Z partner and former Stanford instructor, Nick McKeown's former student. He pushes the heterogeneity debate (agent-written kernels making diverse hardware easier to program) and asks the closing question about rebuilding VMware for agents.
Ron Smith
The Intel recruiter who interviewed Gelsinger at 18 and wrote "smart, aggressive, arrogant. He'll fit right in" 00:03:50.
Ed McCluskey
Stanford professor, "sort of the father" of built-in self-test, with whom Gelsinger argued over practicality while applying BIST to the 386 00:04:59.
John Hennessy
Gelsinger's Stanford thesis advisor, later Stanford president. "John Hennessey was my thesis advisor, and obviously he did okay" 00:05:27.
Alberto Sangiovanni-Vincentelli
Berkeley professor whose group Intel worked with on the first automated placement, routing and timing management for the 486: "we worked with Alberto Sanjivangi, Vincentelli at Berkeley and some of his students for the first, you know, placement, the first routing, the first automated timing management" 00:07:14.
Nick McKeown
Stanford professor and Guido's thesis advisor, credited with the framing that general networks assume unknown packet destinations while AI flows are "large and predictable": "he describes it, well, we built networks to be able to handle any packet going anywhere with no knowledge of where it might go" 00:38:39.
Dario Amodei
Anthropic CEO, referenced through his "constitution" concept as a model for setting policies and guardrails on agents 00:51:39.
Thomas Edison
Invoked by Gelsinger for "the revenge of Edison is upon us" with 800V DC data centers 00:47:00.
5. Operating Insights
Design-to-Silicon Cycle Time Is the Metric That Matters
Gelsinger's example is a good operating lens: if AI cuts your design time to three months but fab, packaging and rack integration still take nine months, workload assumptions go stale before shipping. His target is to compress the back half to "a month or two" 00:11:29. For any hardware company, measure and attack total idea-to-deployed-at-scale latency, not the step AI just accelerated.
Pair Hardware with Workload-Specific Software Investment to Land a Hyperscaler
Gelsinger notes that big buyers choose chips where "it isn't just hardware. It's how do the hardware and software co-evolve" and that these platforms "require investment in the software and the workload evolution to take advantage" 00:19:09. Chip startups should engineer the co-development relationship with one anchor customer early, since that relationship determines whether they survive consolidation.
Design Your Product for Agents as the Primary User, Then Give Humans the Policy Layer
Guido describes the shift when selling infrastructure to agents ("you want the agent to make the pick for your VM offering. It changes a lot how much complexity you can have... Humans are a lot more patient than agents" 00:50:49). Gelsinger's synthesis is two-sided: the core abstraction serves agents (security, performance, migration), while humans get "the policies, describing the constitutions... the dashboards" 00:52:07. The operator takeaway is to build separate surfaces for agent consumption and human governance.
Use the "Stay-Layered" Rule When Consolidating Heterogeneity
Gelsinger's comment on how incumbents absorb specialization ("that's what layering, that's what abstraction has always done" 00:21:49) suggests a product strategy: if you are building a specialized component, make it composable into a large vendor's abstraction layer rather than competing with the full stack.
Underwrite Capex Against Power Availability First
Gelsinger's warning ("do I really have the energy to underwrite the capital commitments that I'm making, for cement, for data center racks and GPU purchases?" 00:45:37) is a practical diligence checklist item: confirm interconnection, turbine and generation timelines before committing to GPU and facility spend.
6. Overlooked Insights
Agents Eliminate the Fleet-Management Argument Against Heterogeneity
In an offhand exchange, Gelsinger lists the operational costs of running a heterogeneous fleet (managing, upgrading, connecting, fault domains), and Guido answers "The agents will do all the upgrades. So, that's no longer a problem... Everything gets easy" 00:42:32. This is easy to miss but important: one of the long-standing structural arguments for hardware standardization (operational complexity) is being weakened by agentic operations, which would favor more heterogeneity than Gelsinger's own consolidation thesis assumes. Combined with Guido's point about agent swarms writing optimized kernels overnight, the software-complexity and ops-complexity barriers to diverse hardware are both eroding, leaving capital and power as the real limits.
Power Guard-Banding Wastes Roughly 40% of the Power Envelope
Dropped in passing during the bottleneck list is the claim that "most of the power simulation aspects today are bad... I'm essentially guard banding almost my full power envelope... essentially, I'm losing 40% of my power in guard banding" 00:12:58. Because power is the binding constraint on the whole AI buildout, a 40% recovery through better 3D power and hotspot modeling and tighter voltage regulation would be equivalent to adding massive generation capacity at the chip level. It points to an under-served software/EDA opportunity, power-aware modeling and regulation, that sits upstream of all the grid and turbine discussion.