Google DeepMind's Logan Kilpatrick: Why the Model Eats the Harness
- 01The Agent Harness as Google's New Unifying Layer
- 02The Model Eats the Scaffolding
- 03Coding Agents as a Narrow Superintelligence Moment
- 04Jagged Superintelligence Before General AGI
- 05Agentic AI Is Positive Sum, Not Cannibalistic
- 06Personal Software as a New Mass Market
1. Key Themes
The Agent Harness as Google's New Unifying Layer
Just as Gemini became the through-line connecting all Google products, the agent harness (called "Agentspace") is now becoming the next unifying layer. Logan describes how this same harness powers not just the coding IDE but also Search, the Gemini app, Cloud, and AI Studio — all specializing from a shared base.
"Anti-gravity will be powering a bunch of agent stuff in search, in the Gemini app, across cloud and in AI Studio, which is really exciting." 00:03:39
The Model Eats the Scaffolding — Harnesses Have a Shelf Life
Logan makes a pointed argument that the current alpha in agent harnesses is temporary. As models improve, they will internalize the scaffolding, making custom harnesses less differentiated. The cycle repeats: scaffolding leads, model absorbs, alpha moves elsewhere.
"The scaffolding is like oftentimes a couple of steps ahead of like where what is like baked directly into the model. And then what ends up happening is like the model eats that scaffolding and it becomes part of like the native model system... Maybe the agent harness is like the quintessential example of this right now where like everyone's like, ah, we got to go build a harness. And like the harness is where the alpha is. And like, I think that perhaps won't be true, at least in the way that we think of the harness today, in 12 months." 00:39:33
Coding Agents as a Narrow Superintelligence Moment
Logan argues we may already be at a point that qualifies as narrow superintelligence in the coding domain, and that this is enormously consequential even before general AGI arrives. The practical impact: individuals can now pursue ideas that previously felt out of reach.
"I feel like I can tackle more ambitious problems. I feel like I used to kick around ideas and they were like slightly out of reach. And I would just be like, ah, wouldn't it be nice. And now I have the opposite problem, which is I'm kicking around an idea and I'm like, I could probably make this even more ambitious." 00:21:51
Jagged Superintelligence Before General AGI
Logan introduces the idea that we won't get a clean general superintelligence, but rather a patchwork of domain-specific superintelligences arriving in sequence — weighted toward verifiable domains first.
"It feels like we're going to get a bunch of those before we've like solved... it's almost like jagged, like jagged super intelligence, I think is what we'll end up with." 00:22:49
"Things that have like better verifiability obviously are like the ones where you'll see the gains happen more quickly. So like things with like math and finance — actually like science could be a really interesting one." 00:23:01
Agentic AI Is Positive Sum, Not Cannibalistic
Contrary to the fear that AI agents reduce human engagement with Google products, Logan argues the pattern from search suggests the opposite: both human usage and agent-driven usage grow simultaneously, expanding the overall ecosystem.
"At the beginning of the sort of current AI era, like everyone assumed that AI being able to answer questions for you was going to be like negative sum for search. And actually what's ended up happening is it's been incredibly positive sum for search. Like people are searching more, people are doing more... Agents are doing more, at the same time that humans are also searching more." 00:05:43
Personal Software as a New Mass Market
The 350,000 Android apps built in AI Studio within a week signals the emergence of software built by individuals for their own use cases — a category that essentially didn't exist before.
"350,000 apps that like probably no one was going to build before. A lot of these are personal too. And so this is where I think... the idea of you building software to solve your personal problem is like very real right now." 00:36:43
Gemini's Post-Training Gains Signal a New Capability Lever
Logan reveals that Gemini 3.5 Flash's improvements over prior Pro models came entirely from post-training, not pre-training — a significant indicator that Google's post-training team has cracked something important and that model progress is no longer gated solely on compute and pre-training runs.
"3.5 Flash was like all post-training gains, which is really cool. So a huge, huge testament to the team, the work that that team did to actually like make the level of gains and like surpass the previous pro model literally just with post-training, which is awesome." 00:17:56
Omni as a True Single-Model Architecture
Google's Gemini Omni is not a routing layer on top of specialized models — it's a single model trained to handle every modality. This is architecturally distinct from prior multi-model pipelines and is the foundation for world-model-level understanding.
"It is a single model, which I think is the important part. Like this was actually part of the original desire was like you were training like eight different models to do all of those things... This is like a true Omni model." 00:31:17
2. Contrarian Perspectives
Startups Have More Opportunity Now Than Before AI — Not Less
The conventional fear is that frontier model labs will eat the application layer. Logan argues the opposite: focus is the startup superpower, and the AI primitives (coding, agents) are closing the gap between startups and incumbents with large codebases.
"One of the outcomes is there's less opportunity for startups in the future. That feels like so far in a way, not what has ended up playing out, which is really positive. If anything, it feels like there's just even more opportunity than there was." 00:43:06
The Agent Harness Is Not a Durable Moat
While the entire startup ecosystem is racing to build proprietary agent harnesses as a defensible layer, Logan predicts these will be commoditized directly into the model within 12 months.
"I think the models will have sort of just like digested a bunch of that. It'll be upstreamed into the model and the alpha will be somewhere else now. It won't be in sort of trying to spin your own harness because the model just like does it natively." 00:40:03
Google's Success Metric Should Be Minimizing Eyeball Time, Not Maximizing It
This inverts the traditional platform playbook. Logan explicitly says that maximizing time-on-product is the wrong goal for the agentic era — success is completing tasks so users can go live their lives.
"Success for Google probably doesn't look like, you know, maximizing eyeball time in front of our products. It's like maximizing outcome for customers to like do the thing that they want to do so that they can go and live their life and do what they want." 00:06:58
Pre-Training Run Timing Explains Competitive Position Better Than Public Benchmarks
Logan suggests that what looks like a Google capability gap to outside observers often simply reflects where large pre-training runs are in their cycle — a non-public variable that makes external competitive narratives misleading.
"You miss all the context of like where the big runs are and where the large pre-training runs are. So it might look from an external perspective that like, oh, you're super behind in some way. And like actually you miss all the context." 00:17:27
World Models and Video Models Are Converging Into the Same Thing
The conventional view treats world models (action-conditioned video models) and multimodal language models as categorically different. Logan argues Omni blurs this distinction, and that the future of game-making and interactive simulation will be driven by this converged architecture rather than traditional world models.
"I think the definition of world models will blur... I do feel like this like world model video model thing is going to change and play out in a different way than was obvious before." 00:28:17
3. Companies Identified
Google DeepMind / Gemini
Google's AI research and product division. Discussed extensively as the engine room of Google, powering 13 billion-user products with Gemini models, launching Agentspace, Omni, and leading agentic AI development.
"Deploying Gemini to billion user products is a problem that like only two companies in the world have. We have 13 of those products." 00:48:00
Agentspace (Google / formerly tied to Windsurf acquisition)
Google's agent harness / coding IDE ecosystem. Built on the back of the Windsurf deal and talent acquisition, it powers coding, consumer agents, and enterprise agentic experiences across Google products.
"Anti-gravity is a lot of things... You have sort of a core IDE. You have sort of like the agent first experience if you want it on the web. You have a CLI. You have an SDK." 00:03:11
Open Router
An API aggregator for LLMs that tracks cross-model token consumption. Cited as a key data source for measuring how much AI intelligence is actually being deployed in the world over time.
"Something like Open Router, for example, is like measuring, you know, the total token consumption that's happening. And so you can sort of like see these trends play out over time of like how much more intelligence is in the world, you know, now versus a year ago." 00:12:17
Anthropic
AI safety-focused model lab. Mentioned for its distinctive culture rooted in Dario Amodei's personality — described as somewhat esoteric and intellectually rigorous.
"I think Anthropic is a very interesting place... He seems like an interesting guy. And so, uh, somewhat esoteric. And so that seems like they're sort of like that in the DNA and the culture of the company." 00:45:59
Kaggle
Google's AI benchmarking and data science community platform, now working with DeepMind to build a "game arena" for testing model capabilities in games.
"Our team actually in Kaggle, which is sort of a bunch of the AI benchmarking stuff we do in GDM, sort of works with GDM to build this, uh, game arena, which is our way of sort of like testing." 00:26:44
4. People Identified
Demis Hassabis
CEO and co-founder of Google DeepMind, Nobel Prize scientist. Described as having a deeply scientific orientation toward AI that permeates DeepMind's culture. His original mission to solve disease and science grounds the lab's purpose beyond competitive benchmarking.
"Demis is a Nobel prize scientist. And like the sort of OG of a lot of this stuff and sort of, you feel that in the DeepMind culture." 00:45:32
Tulsi (leads the model team at Google DeepMind)
Head of the model team at DeepMind. Named as exceptional by Logan, who specifically called her out for potential podcast appearances.
"I was giving a talk and was on stage with my friend Tulsi, who leads the model team, who I don't know if you've ever had on before, but she's amazing. I love Tulsi." 00:00:03
Noam Shazeer
Legendary AI researcher, co-inventor of the transformer architecture, returned to Google. His return is cited as a signal of the exceptional talent concentration at DeepMind right now.
"I've heard Sergei's back. You guys have Noam Shazeer back." 00:43:53 (attributed to Sonya Huang)
Sergey Brin
Google co-founder, noted to have returned to active involvement at Google during this AI moment.
"I've heard Sergei's back." 00:44:12 (attributed to Sonya Huang)
Josh (leads Gemini app / Mac OS app)
Referenced as a peer of Logan's at Google, leading the Gemini consumer app team and cited specifically for building a Mac OS app faster than any team in Google's history using agentic coding.
"Josh's team did this with the Gemini Mac OS app and sort of like end to end delivered an app sort of faster than any team had ever delivered a Mac app at Google. And it's because of agentic coding." 00:19:59
5. Operating Insights
Build the Product to Build the Model — You Cannot Separate Them
Logan reveals a key lesson about why Google's coding model lagged: you cannot make a great coding model for long-running software engineering tasks without actually having a product that does that work at scale. The data flywheel from the product is what trains the model. This is a template for any AI company: the product and the model must co-evolve.
"It's actually really hard to make a great coding model for this like really long running SWE work if you don't actually have a product that does that. And so I think like Google realized that. That's why the sort of like Windsurf deal happened." 00:15:55
Use Competitor Products Deliberately — It's How You Understand the Ecosystem
Logan advocates explicitly for employees using all competing models, not just their own. The reasoning is epistemic: you can't understand your own product's gaps if you only live inside your own ecosystem.
"It's so healthy to be using other models just because like it's sometimes hard to like actually grok what's happening in the ecosystem if you're not... I use all the models, I use all the products." 00:18:23
Dogfooding at Scale Is a Structural Competitive Advantage
Having 100,000+ engineers using and giving feedback on your model is a compounding moat that most companies cannot replicate. Logan frames this as something Google should be more deliberately leveraging.
"DeepMind has and Google more broadly has like 100,000 plus incredible engineers who are using the models and giving feedback. And like it should be a competitive advantage for Google because we have that scale of sort of engineering resources and like the depth of the talent and can run, you know, A-B tests and live experiments and all that stuff." 00:18:44
Reset Your Ambition Level When Using AI Coding Tools
For operators and founders, Logan describes a personal reframe: AI coding shouldn't just make existing ideas faster — it should push you to make the idea itself more ambitious. The risk is defaulting to the MVP when the tools now support something much larger.
"I have the opposite problem, which is I'm kicking around an idea and I'm like, I could probably make this even more ambitious. And sort of, it adds a different layer of sort of, um, responsibility or like some, a different layer of burden actually, because I'm like, oh, I can't just like do the sort of MVP of this. Like I actually need to like go 10 steps further because the technology enables me." 00:21:51
6. Overlooked Insights
"Harness Bench" — A Missing Benchmark That Would Reshape Model Selection
In a single throwaway sentence, Logan identifies a benchmark that doesn't exist but should: a test of how well each model adapts to arbitrary agent harnesses, not just its own provider's harness. This matters enormously because enterprise buyers are choosing models partly on the assumption that proprietary harnesses lock them in. If a neutral benchmark showed that certain models are universally harness-agnostic, it would fundamentally change procurement decisions — and likely favor models that generalize best.
"We need something like harness bench, which is like actually measuring like how good are all these different models at adapting to all the different harnesses. I feel like that seems like a reasonable thing we should measure as an ecosystem. And I'd be curious to see like what models are actually best." 00:40:45
This is a concrete product/company opportunity: whoever builds and publishes "harness bench" credibly becomes the standard-setter for agentic model evaluation — a highly leveraged position.
Crypto Is the Dominant Use Case Driving AI Studio Developer Activity
In a brief aside, Logan reveals that the largest category of apps being built in AI Studio — approximately 20% — is finance, and specifically crypto-related. This received zero follow-up in the conversation but is a significant signal: the crypto developer community is disproportionately early and aggressive in adopting AI coding tools, and the intersection of crypto and AI agents (autonomous on-chain agents, crypto portfolio tools, DeFi automation) is a real and already-active developer market, not a theoretical future one.
"I think it was like it's like 20%, like finance related stuff... I think it's something around crypto actually, I think is what people are doing a lot of stuff with." 00:26:13