Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/THE A16Z SHOW/AI Can Write Code. Why Isn’t Sof…
POD
// EPISODE
THE A16Z SHOW

AI Can Write Code. Why Isn’t Software Better?

DATE September 28, 2026SOURCE THE A16Z SHOWPARTICIPANTS BEN HOROWITZ, DIOGO ALMEIDA, MARTIN CASADO
// KEY TAKEAWAYS6 ITEMS
  1. 01The Automation Paradox: Extreme Intelligence, Near-Zero Deployment
  2. 02Coding Agents Make the Same Software Faster
  3. 03A New Primitive: Natural Language In, State Machine Out
  4. 04Reliability, Not Demos, Is the Real Product
  5. 05RLHF Generalization Was the Founding "Aha," and Its Failure to Become AGI Was the Founding Disappointment
  6. 06Human Evaluation Has Optimized AI for the Wrong Target

1. Key Themes

The Automation Paradox: Extreme Intelligence, Near-Zero Deployment

The episode's animating question is why, despite AI being demonstrably brilliant, so little real-world work has actually been automated. Diogo Almeida frames this as almost morally urgent. "Where the fuck is all the automation? Like this is like so unbelievably tragic... AI is so unbelievably smart. And yet, so not that I hate on chatbots or coding agents. I love them myself. But it's like so useless at all other stuff. And it's tragic." [00:00:00] He extends this to the industry's biggest lab: "OpenAI has been trying to automate customer service since 2020." [00:00:14]

Coding Agents Make the Same Software Faster — They Don't Make Software Better

Martin Casado draws a sharp distinction between accelerating existing software patterns versus expanding what software can fundamentally do. "If you use something like Claude Code or Codex, which is great, or Cursor, which is great, they write code. But that code is the same thing a human being would have write. Maybe it's better, maybe it's worse, but it's basically still code, just like code looked 10 years ago." [00:04:57] He adds a stark empirical point: "it doesn't matter how much AI coding agents you use, the software actually isn't getting better. Maybe you're writing it faster. It's, like, arguably getting worse just because, like, there's less oversight." [00:00:07]

A New Primitive: Natural Language In, State Machine Out

JEV's core innovation, per Casado, is a new programmable primitive that any code (human- or AI-written) can call to make probabilistic, natural-language-driven decisions inside software rather than around it. "The thing with JEV is whether or not you're Claude Code or a human, you have this new primitive, this new thing that you stick in your code that actually expands the power of software." [00:05:15] Diogo frames the same idea from the builder's side: "I want to expand what software itself can do such that things that should be automatable can then be automatable." [00:00:14] Notably, the interface is deliberately named "input state" — "it's meant to be the inside of programs." [00:08:08]

Reliability, Not Demos, Is the Real Product

Diogo repeatedly stresses that the AI industry's obsession with flashy demos is an anti-pattern he's explicitly avoiding. "If we were going to be really intellectually honest and we are really aiming for the North Star of automation, we cannot fall into the same anti-patterns that AI has fallen into, which is really focusing on outliers and demos." [00:15:57] He wants automation invisible and trustworthy: "I want them to work in the background such that, like, someone would trust that to run and not page them." [00:15:57] He defines reliability with unusual precision — distinguishing uptime, determinism, and a third category he calls robustness: "similar intelligence every time... it doesn't have to be the similar function every time. But it needs to be smart every time." [00:27:03]

RLHF Generalization Was the Founding "Aha," and Its Failure to Become AGI Was the Founding Disappointment

Diogo traces his conviction back to OpenAI's early RLHF results around Q4 2021. "My favorite query was, why is it important to eat socks before meditating?... the models were able to, like, make plausible human-looking answers for this. And that, to us in the team, was the thing that clicked, like, this is not cheating." [00:18:22] But the subsequent letdown reshaped his thinking: "I really thought that that model had, like, a decent chance of being AGI. And when it didn't, that was, like, when my whole world came crashing down." [00:18:51]

Human Evaluation Has Optimized AI for the Wrong Target

A central critique: because humans judge model quality, labs have optimized for looking impressive to human evaluators rather than for actually automating tasks. "I think GPT-3 was actually quite calibrated back in that day. But because humans evaluate how good the models are, it looks really good because they are the judge. But we've been optimizing that judge instead of the automation part. And that has been the missing thing." [00:20:41]

SaaS Isn't Dying — It's About to Get Dramatically More Valuable ("Inverse SaaSpocalypse")

Against the prevailing narrative that AI will gut SaaS incumbents, Diogo argues incumbents with distribution and workflow knowledge are best positioned to benefit. "I think SaaS will be one of the largest winners of, like, the whole AI game... I think that they are the best positioned to know what workflows to automate." [00:31:02] Ben Horowitz builds on this: "so much of a SaaS company's capital investment is actually getting to all the customers. And so if you've gotten to all the customers and then you make... the software, like, way, way better. That's a hell of a thing." [00:32:05]

An Entire Era of Probabilistic Programming Is Reopening

Diogo connects JEV to a decades-old, previously abandoned research tradition. "There's a huge history of probabilistic programming that basically died in, like, the 70s... you could also call JEV, like, neuropsychology AI." [00:36:02] He sees this eventually penetrating deep infrastructure, not just user-facing apps — analogized to TCP/UDP-level plumbing: "when I think of AI... I work backwards from AI-based economic revolution... what percentage of the calls to AI... are, like, for human consumption... And I think that it's going to be many nines in the guts, but it will start at the first layer." [00:38:44]

Coding Agents Are Good at Syntax, Bad at Architecture

Diogo's experience using coding agents surfaces a specific, actionable capability gap. "My experience is that they are really good at syntax and really— They're bad at semantics. I would say incredibly bad at architecture... to me, architecture is, like, the most human creative part of software." [00:28:48] He frames adoption as an ROI/speed tradeoff: "sometimes speed is the knob for your company or project to turn. Like, you're willing to do a 50th percentile architecture instead of a 60th because you want to move faster." [00:29:43]

2. Contrarian Perspectives

AGI-as-defined-by-OpenAI Is Achievable, But We're Not on the Path to Recursive Self-Improvement

Diogo, a former OpenAI researcher who worked on early GPT capabilities, explicitly breaks from the "one brain to rule them all" narrative that pervades AI discourse. "For nuanced reasons, I don't think we are on the path of RSI. And I still don't think we're in the path of RSI... I do think that what OpenAI defined as AGI is extremely doable. Automating most of the world's economically valuable work actually sounds like... a lot of it is very rote and simple." [00:00:44] This is a direct rebuke of "mono model Kool-Aid": "will that one, is that one brain really on the path to rule us all? Like we have not automated really basic things that I don't think we want people to be doing." [00:13:55]

The AI Industry's Doom-and-Gloom Culture Is a Misread of What's Actually Happening

Rather than covering up or fearing AI's implications, Diogo argues most of the industry simply "doesn't get" developers and is having the wrong emotional reaction entirely. Ben Horowitz highlights this as rare: "if we had any other kind of like big lab leader, even if they had joy, they would cover it up. And then your view is so different. You're like, no, we're going to create a way better world." [00:00:27] Diogo: "it's like an ML level concern while everyone else is having like a JEV party... I think if you don't like get developers, it'll be hard to understand what's really going on." [00:13:27]

Biologically-Inspired AI Approaches "Never Worked" — Pragmatism Beats Mimicking the Brain

Diogo dismisses a foundational narrative that neural nets succeeded because they mimicked biology. "I'm not a fan of, like, biologically inspired stuff at all... I think it's never worked. It's useful to motivate crazy people to work on things for decades until it works and then they refine it into, like, the engineering version." [00:36:35] He goes further to debunk the folklore: "a lot of the stories about how it worked were not accurate." [00:37:02]

The "Data Distribution" Excuse for Failed Real-World Automation Is Overstated

When Martin Casado offers the common explanation that automation hasn't reached the physical/messy world because we lack training data for that distribution, Diogo pushes back. "I don't entirely buy the data argument, in my opinion... I don't think that in my, like, canary in the coal mine situation, we need to automate that long tail... it should be an ROI decision for people who, like, automate stuff." [00:22:28] His point: the bottleneck isn't data scarcity for edge cases, it's that we haven't even automated the obviously rote, high-volume, easy stuff.

Coding Agent Value May Actually Shrink as Smarter Primitives Emerge

Martin Casado floats a genuinely uncomfortable idea for the coding-agent boom: that tools like Codex/Claude Code could become less valuable once primitives like JEV exist, because the real bottleneck isn't code generation but decision-making logic. "It occurs to me that actually the value of things like coding agents goes down if you have a primitive like this... Codex builds all the software for me, but it doesn't actually use JEV. And so, like, the software itself creates it somewhat limited." [00:28:09]

3. Companies Identified

TypeSafe AI — Diogo Almeida's company, creator of JEV, described as building "AI for software" that gives developers a new primitive for expanding software's native capabilities rather than just generating more of the same code faster. Mentioned throughout as the episode's central subject and, per Ben Horowitz, "a whole movement towards a positive future." [00:13:12] "TypeSafe is making AI for software. You know, we want to make AI powerful, not just for humans in the loop, but to actually build real software." [00:02:50]

OpenAI — Where Diogo worked and helped ship early RLHF/GPT capabilities; also cited as an example of overpromise/underdeliver on automation despite years of trying. "OpenAI has been trying to automate customer service since 2020." [00:00:14]

Google Brain — Mentioned as a stop in Diogo's career history between an earlier startup and his eventual return to AI at OpenAI.

Anthropic (Claude Code) and Codex / OpenAI Codex and Cursor — Referenced repeatedly as best-in-class coding agents, praised for speed and quality but critiqued for only accelerating "the same kind of software we already have." [SPEAKER_00, 00:00:07] Gary Tan's description of them as "just in time software" is cited approvingly by Diogo: "Incredible way to describe what they're doing. Like it makes software on the fly." [00:03:39]

SaaS incumbents (broadly) — Framed as the biggest coming winners of the AI wave due to distribution and workflow knowledge, contrary to "SaaSpocalypse" fears. "I think SaaS will be one of the largest winners of, like, the whole AI game." [00:31:02]

a16z Growth Fund's internal tooling ("Muse") — Mentioned by Ben Horowitz via a conversation with David George as a real, if mundane, automation win: "he said, I finally canceled my New York Times subscription." [00:15:37]

4. People Identified

Diogo Almeida — Founder/CEO of TypeSafe AI, creator of JEV; former OpenAI researcher who worked on early RLHF and GPT capabilities, later at Google Brain, and earlier at a startup with Jeremy Howard. Former competitive "mathlete" and Kaggle competition winner. Described by Ben Horowitz as "a bit of a hero to both Martin and me." [00:01:44] His personal design philosophy: "My brand is pragmatism. Incredible pragmatism." [00:36:35]

Isabel Guillon — Co-inventor of the SVM (support vector machine), Kaggle competition host who mentored Diogo early in his career. "She just basically saw that I was like this person who really didn't fit into the research community and then adopted me and showed me like it got me to meet all the AI people." [00:11:21]

Jeremy Howard — AI educator/entrepreneur (fast.ai) with whom Diogo worked at an early startup before joining Google Brain and later OpenAI. "It was like a startup with Jeremy Howard... I love Jeremy." [00:11:47]

Eric (Diogo's TypeSafe co-founder) — Noted by Ben Horowitz as coming from a Bayesian/probabilistic and biology background, connecting to JEV's probabilistic-programming roots. "So your co-founder, Eric, came from that background... They kind of Bayesian... He did a lot of biology." [00:36:28]

Gary Tan — Credited by Diogo with the phrase "just in time software" to describe coding agents like Claude Code and Codex, called "an incredible way to describe what they're doing." [00:03:39]

David George — Runs a16z's growth fund; mentioned by Ben Horowitz in an anecdote about using "Muse" to automate reading tasks and cancel a New York Times subscription. [00:15:09]

Ilya (Sutskever) — Referenced via an internal OpenAI cultural in-joke about the vagueness of "AGI": "people used to describe it as Ilya in every if statement." [00:19:36]

5. Operating Insights

Design Your Interface Names to Encode Your Product Philosophy

Diogo reveals that JEV's API deliberately calls its input "state" rather than something more generic, as a subtle but intentional signal of the product's positioning as infrastructure inside programs, not a chat layer alongside them. "Even like our interface, like calling the input state, this is intentional... It's meant to be the inside of programs." [00:08:08] This is a transferable tactic: naming choices in developer-facing products can embed and reinforce strategic intent.

Optimize for "Intelligence Per Dollar," Not "Intelligence Per Second," as a Product North Star

Diogo explicitly names the metric he uses to make architecture tradeoffs, distinguishing it from the industry's usual obsession with raw capability or latency. "Intelligence per dollar is my North Star right now. And it could be wrong, just to be clear. Intelligence per second might be more valuable in the short term." [00:08:08] This is a concrete, quotable framework other infra/AI builders could adopt to structure prioritization debates.

Avoid the Demo Trap — Build for Invisible, Unattended Trust, Not Showcase Moments

Diogo's team intentionally resists highlighting flashy use cases because the entire value proposition is things running reliably without human oversight. "A lot of people ask me, like, what are your favorite use cases? And I'm like, I'm not sure if they work. I want them to work in the background such that, like, someone would trust that to run and not page them." [00:15:57] For operators, this reframes what "impressive" should mean internally — the boring, invisible win beats the demo win.

Slow Down Launches Deliberately to Protect the Reliability Bar Even Under Pressure to Ship

Diogo states outright that TypeSafe held back releasing capability for a long period specifically to protect the reliability guarantee, even though this cost them speed-to-market advantage. "We could have released so much sooner. I don't think people realize that." [00:40:55] Combined with: "the amount I care about reliability is, it's a lot. Like, reliability is what this thing is. If you don't understand that, it'll be very hard to make, like, a copycat that's benchmarked." [00:25:36] This is a direct articulation of reliability-as-moat, a strategic reason to intentionally under-ship in a hype-driven market.

Use Adversarial, Off-Distribution Test Queries to Validate Genuine Generalization vs. Memorization/Cheating

Diogo describes the specific technique OpenAI's RLHF team used to convince themselves the model wasn't "cheating" — constructing absurd, verifiably-not-on-the-internet prompts. "My favorite query was, why is it important to eat socks before meditating? We'd made sure that was not on the internet beforehand." [00:18:22] This is a reusable diligence technique for anyone evaluating whether a model (or vendor's model) is actually generalizing versus pattern-matching on training data.

6. Overlooked Insights

The Real PR-Size Data Point Reveals How Little Coding Agents Actually Change

Buried in a single aside, Martin Casado drops a concrete, empirical fact from his time at Google that undercuts much of the coding-agent hype: most software changes in large companies are tiny. "If you actually look at, like, the average PR for a large company, it's like, 10 lines, right? Seriously... I've been at Google. We actually did the study. So it's like 10 lines. So, like, you're automating 10 lines." [00:33:40] This is a significant, underappreciated data point for investors evaluating the "AI writes code" category: even massive gains in code-generation speed may only be optimizing an already-tiny slice of engineering effort (small, incremental PRs), not unlocking transformative new capability — reinforcing why Casado and Diogo both argue the more interesting frontier is expanding what software can do rather than how fast it's typed.

Customer Support "Automation Rates" Are Statistically Misleading — True Coverage Is Much Lower Than Advertised

Martin Casado surfaces a critique of an entire category of AI-support-automation claims that likely applies broadly across the "AI agent" investment landscape. Companies claim to "answer 95% of all help desk calls," but "you actually look at the data. They're all the same. And you realize it's all password resets. And then, like, but if you did it by, like, uniqueness, it was only something like 50% or something." [00:24:09] This is a quietly devastating point for due diligence on any AI-automation startup citing high resolution/automation percentages — investors should ask whether the metric is volume-weighted (dominated by trivial repeat cases) or diversity-weighted (reflecting genuine breadth of automated capability), since the two paint radically different pictures of defensibility and TAM.