Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/NO PRIORS/The Future of Frontier Model Arc…
POD
// EPISODE
NO PRIORS

The Future of Frontier Model Architectures with Walter Goodwin, Fractile Founder and CEO

DATE October 2, 2026SOURCE NO PRIORSPARTICIPANTS SARAH GUO, WALTER GOODWIN
// KEY TAKEAWAYS6 ITEMS
  1. 01Memory bandwidth, not FLOPs, is the binding constraint on frontier inference
  2. 02The gap between SRAM speed and DRAM capacity is the unsolved problem in fast inference
  3. 03Inference economics collapse to cost per gigabyte of memory
  4. 04Speed's real prize is long-running agents, not snappier chatbots
  5. 05Full-stack integration enables an agile, closed-loop chip design cycle
  6. 06Hardware co-designing the model landscape: scaling laws for bandwidth
In this episode

1. Key Themes

Memory bandwidth, not FLOPs, is the binding constraint on frontier inference

Goodwin's core thesis is that chips have scaled compute far faster than the thing LLMs actually crave. He quantifies the imbalance: "We've scaled flops like a million fold in the last 20 years. Memory bandwidth has gone up about 40x in the same time frame." 00:30:45 Autoregressive, low-batch LLM decoding is bandwidth-bound, so "a new LLM desperately wants to have orders of magnitude more bandwidth to memory in order to run faster." 00:16:22 This is the foundation of Fractile's entire product bet.

The gap between SRAM speed and DRAM capacity is the unsolved problem in fast inference

Existing fast-inference chips (Groq, Cerebras) have super-high-bandwidth SRAM but tiny capacity. Goodwin says this creates a "tantalizing mismatch": "they're not running long context attention. You still flip back to a GPU to do that." 00:13:53 Fractile started on SRAM but pivoted: "one of the things that we started to worry about in kind of towards the end of 2023, and certainly in 2024, was the scalability of this approach," because of parameter growth and "this kind of growing context length." 00:11:43 The new platform "unites... the scalability of these kind of higher capacity, lower cost DRAM memories that you have on a GPU, on a TPU, with all of the speed advantages that you get from a Grok chip or a Cerebras chip," ramping in the second half of next year. 00:12:29

Inference economics collapse to cost per gigabyte of memory

A key operating-economics reframe: "when you look at what it looks like to run inference at data center scale for thousands of users, the economics actually collapsed down to essentially a kind of cost per gigabyte of the memory that you're employing." 00:14:36 This is why Fractile is working on getting "aggressively high bandwidth from... the world's lower cost memory, DRAM memory" rather than exotic, expensive memory. 00:14:36

Speed's real prize is long-running agents, not snappier chatbots

Goodwin dismisses the obvious use case: "The snappier chatbot is kind of the faster horses of kind of fast inference." The real unlock: "taking a multi-trillion parameter model and run it comfortably at many thousands of tokens per second, in our view, is a fundamental elevator on capability for AI, is in taking these very long running agents and making them radically faster." 00:13:24 Speed also buys quality: frontier players need "the fastest deployment so you can do the most reasoning in the shortest amount of time." 00:33:08

Full-stack integration enables an agile, closed-loop chip design cycle

Fractile (about 150 people, "skinny in every single one of those sectors") owns front-end design, physical design, back-end implementation and advanced packaging in-house. Goodwin contrasts this with the industry's waterfall handoff: "you get to a certain level, and then you hand off to another partner. And to some extent, you're kind of at the mercy of how that other partner then behaves." 00:09:38 The in-house loop matters because "in the AI chip space more than any other kind of chip play before, you have to really place your bets correctly. So you have to have a lot of skill and a lot of luck. And you also then need to strike very fast." 00:09:09

Hardware co-designing the model landscape: scaling laws for bandwidth

With roughly "25 times more bandwidth per chip than an HBM based chip," Fractile explores "scaling laws for bandwidth." 00:29:24 Examples: MoE models should be far sparser (1 in 16 to 1 in 128 or 1 in 256, saving flops at iso-intelligence), but that "becomes incredibly prohibitive on today's HBM based GPUs, XPUs to serve... You end up serving at really, really low MFU." 00:29:50 Likewise, bandwidth-light attention variants tend to be more flops-hungry. Raising bandwidth lets you conserve flops, "a multiplying factor on global throughput for these models." 00:31:11

Compressed design cycles buy "more shots on goal," not a new chip every few weeks

Goodwin is explicit that physical constraints (fab cycle times of "three to five months, even in a kind of super hot lot scenario," a "three to five year amortization window," data center power) cap how fast chips can turn over. 00:19:08 The value of shortening the cycle is optionality: "a rolling frontier of bets that you hope are deeply aligned in which you're ready to trigger the ramp of." 00:00:00 The prize: "if you can just find a way to structurally carve out a three to six months advantage, you will be winning all of those deployments." 00:00:30

The new bottleneck after AI-accelerated thinking is wall-clock experiments

In both AI research and chip design, smarter reasoning shifts the bottleneck to slow physical or computational steps: "now we're just experiment bottlenecked. We're compute time bottlenecked." 00:25:01 Place-and-route tools "run for days on end," so Goodwin argues for AI surrogate and "fuzzy placement" models to speed iteration while Cadence and Synopsys retain final sign-off. 00:26:30

Market structure: diversity of supply plus a premium speed tier

Goodwin sees every large deployer running as many platforms as possible. First-party chips are architecturally similar to Nvidia's, so their effect is partly price leverage ("a bit of a joke today that the sort of first party efforts, their primary purpose is to reduce the price that people pay NVIDIA"). 00:32:12 Chips that offer new capabilities are different: "I suddenly see for what we're doing a real need for everybody that wants to deploy AI at the frontier." 00:32:39

2. Contrarian Perspectives

Most AI ASICs are "identikit," and the industry underinvests in true full-stack differentiation

While the market sees an "embarrassment of riches" in AI chips, Goodwin sees convergence: "all of these chips, you have HBM... the same kinds of bets on tensor cores to do your matrix multiplications, the same advanced packaging with TSMC." 00:03:52 Because many hyperscaler efforts share the same handful of back-end ASIC houses (Broadcom is "the largest, you know, $2 trillion company"), "there's a relative dearth still of kind of efforts that go all the way up and down the silicon stack and try and build kind of fundamentally new capabilities." 00:04:12

Hyperscaler chip programs are mostly a pricing weapon against Nvidia

Rather than technological moonshots, Goodwin says "the first party efforts, their primary purpose is to reduce the price that people pay NVIDIA... those efforts are kind of architecturally quite similar, right? So it's a bet that is not on enabling a fundamental capability that those other chips cannot enable." 00:32:12

Frontier labs shouldn't (and won't) roll up chip design in-house

The conventional "vertical integration" instinct says labs will build their own silicon. Goodwin argues the opposite is structurally rational: "Suppose I'm lab one and I've gone all in on some proprietary silicon. And then lab two discovers some new computational breakthrough... that only works on the chip that they've decided to deploy. I could die in the nine months before I get to deploy enough of that chip." 00:34:27 He concludes that "going all in on those differentiated bets at the chip player is sort of an irrational and a very dangerous move for anybody that is playing at the frontier." 00:34:50 Hence "over multiple decades, the world sustains some kind of frontier third party chip players." 00:33:57

Don't expect a new chip every few weeks, even with AI-accelerated design

Contrary to the "everything compresses to software speed" narrative, Goodwin says: "I'm not a believer in the idea that we will ultimately get to a place where we are shipping a fundamentally new chip, you know, every few weeks just because we've shortened that down. Because I think this is a physical world." 00:20:26 Financing, power and install latency impose real limits.

Divide expert timelines by four

When Sarah relays a top-three semiconductor CEO's estimate of 10 years until architect intent maps to a usable GDS2 file, Goodwin responds: "a heuristic that has been successful over the past couple of years is to just always question your sort of logical assumption and then divide it by four on timescale... And I probably subtract a bit from that." 00:23:27

3. Companies Identified

Fractile

Founded in summer 2022 by Walter Goodwin, about 150 people, building very fast inference chips for the largest models, with a platform ramping in H2 next year. Mentioned as the subject company. "Fractile is a chip company. We build very, very fast inference chips for the world's largest models. This sort of bet on speed is something which we've had right from the outset." 00:01:24 Differentiator: "25 times more bandwidth per chip than an HBM based chip." 00:28:55

Nvidia

GPU market leader. Mentioned as the benchmark competitor with an internal chip portfolio Fractile aspires to match: "if you crack open an NVIDIA system, it has anywhere between kind of six and nine custom chips all built by NVIDIA to come together to build something that is really, really potent." 00:00:00

Groq

SRAM-based fast-inference chip company. Mentioned as a reference for speed and as the architecture Fractile initially resembled. "We were working on a little like Grok or Cerebrus, working on an SRAM based chip." 00:10:34 Goodwin notes SRAM "can drive you to thousands of tokens per second on these language models." 00:10:58

Cerebras

Wafer-scale SRAM-heavy inference chip company. Mentioned alongside Groq for "all of the speed advantages that you get from a Grok chip or a Cerebrus chip." 00:12:29

Google (TPU)

Pioneer of the proprietary AI ASIC. "Google kind of kicked this off with the TPU more than 10 years ago," and the tensor core is an example of an architect-led insight into workload dominance by matrix multiplication. 00:02:46

Broadcom

Largest of the "ASIC houses" that deliver hyperscaler custom chips, described as "you know, $2 trillion company that help people kind of realize these designs." 00:03:32 Owns analog IP for chip-to-chip connectivity and physical placement for specific process nodes. 00:06:50

TSMC

Foundry at the center of the supply chain; chips ship as a GDS2 "bitmap," and advanced packaging is shared across the industry's AI chips. 00:06:32

Meta (MTIA), Microsoft (Maia), OpenAI ("Jalapeno")

Hyperscaler/lab proprietary chip efforts developed with ASIC-house partners. 00:03:09 Mentioned as examples of the similarity across first-party platforms.

AMD

Cited as both a GPU competitor sharing HBM architecture and as a lever buyers use: "can I buy something from AMD so that maybe I then get a better price from NVIDIA." 00:32:12

Cadence and Synopsys

EDA tool vendors. Goodwin calls their work "immensely valuable": "the final checkmark that says this thing is what we call DRC and LVS clean. It conforms to the rules that that foundry has set." 00:26:30 Final sign-off is not expected to change "for quite some time."

Kimi (Moonshot open-source models)

Mentioned as the competitive pressure from behind for frontier labs: "otherwise we'd all be using Kimi models all the time." 00:33:08

Chinese open-source model labs (generally)

Cited as the churn frontier for attention mechanisms and MoE sparsity: "the exact nature of an attention mechanism will churn every couple of weeks in terms of what's kind of the frontier." 00:17:03

4. People Identified

Walter Goodwin

Founder and CEO of Fractile. Notable for the early (2022) bet on inference speed and for a pivot from SRAM to high-bandwidth DRAM-based architectures. "The bet on speed, that was something that kind of really informed a lot of what we then did as a kind of architectural response." 00:10:34

Sarah Guo

Host of No Priors. Notable for relaying a top-three semiconductor CEO's off-record 10-year estimate for architect-intent-to-GDS2 automation. 00:23:02

Henry Ford

Cited via the apocryphal "faster horses" line, used by Goodwin to frame chatbot speed as an unimaginative use of fast inference. 00:12:57

Baron Millidge

Mentioned (transcribed as "Baron Milledge") for appearing on Dwarkesh Patel's podcast with the idea that "perhaps you should think for 100 years... before you fire off an experiment and then think for another 100 years of human equivalent with these models about the results." 00:27:40

5. Operating Insights

Place bets ahead of the workload, and advocate to customers

Rather than reacting to models, Fractile tries "to run ahead in many places and look at where would we change a model architecture to be more exceptionally aligned with our bet? What do we think is the scaling law vector for the particular bets that we are taking? And then we can advocate to our customers and our partners on that." 00:07:57 This builds the "muscle" for designing chips that won't ship in volume for one to two years. 00:08:13

Stay "skinny in every sector" to keep a closed loop

Rather than building deep in one function, Fractile runs a ~150-person team with thin coverage across front-end, physical design, back-end and packaging. 00:08:47 The payoff is a "much more agile closed loop," which matters because the cadence of this industry is set by workloads that "churn at the pace of essentially software." 00:17:31

Remember Amdahl's law for organizational transformation

Adopting AI in one stage of a long chain yields little if the rest stays conventional: "there isn't necessarily an overwhelming need to or return on radically changing your processes as a front-end focused chip design startup, adopting AI... to shorten your timelines from being maybe 12 months of front-end design work to a tiny fraction of that. If there's then going to be kind of the conventional bottlenecks." 00:18:13 The lesson for operators: owning the end-to-end chain is what lets you reorganize around AI productivity.

Convert AI productivity into more parallel bets, not headcount cuts

Goodwin's framing of the AI-productivity choice: "One is, oh no, we're not going to have as much to do. Maybe there's job loss. The other is, we get to do more things." His aim is to build "our own responses" to Nvidia's multi-chip system rather than one chip. 00:21:46

Invest in the surrogate/approximation layer around expensive bottlenecks

When thinking gets cheap, the slow deterministic step becomes the constraint. Fractile looks at "the guts of some of those algorithms" for "a crude approximate way that you could do some of this," while leaving final sign-off to incumbents. 00:26:30 Spend AI effort on accelerating the trial loop, not just the ideation.

6. Overlooked Insights

AI-accelerated thinking converts into a demand for more reasoning tokens before each physical action

Mentioned almost in passing, but it implies a durable workload shape: as experiments (chip placement, AI training runs, wet-lab) stay wall-clock-expensive, rational actors will front-load enormous amounts of inference. Goodwin: "it becomes absolutely, it's like almost, you know, a moral compunction to think harder before every single experiment that you fire off." 00:27:20 And: "We're going to be generating a lot of reasoning tokens before we go and place a given circuit." 00:28:09 This is the demand-side logic for ultra-fast, large-model inference: the more expensive the physical step, the more valuable high-speed reasoning becomes.

Today's fast-inference chips fall back to GPUs for long-context attention, a hidden structural weakness

Said as a technical aside, it undermines the narrative that SRAM-based speed chips are full substitutes for GPUs: "if you look at the sort of technical detail of how they actually get employed and deployed today, they're not running long context attention. You still flip back to a GPU to do that." 00:13:53 As agent contexts grow, this gap becomes the opening for any architecture that pairs high bandwidth with economical capacity.