Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/THE A16Z SHOW/Aaron Levie, Steven Sinofsky & M…
POD
// EPISODE
THE A16Z SHOW

Aaron Levie, Steven Sinofsky & Martin Casado: How Do You Secure a World of AI Agents?

DATE September 26, 2026SOURCE THE A16Z SHOWPARTICIPANTS AARON LEVIE, ERIK TORENBERG, MARTIN CASADO, STEVEN SINOFSKY

1. Key Themes

The "Pacing" Debate Is Internally Incoherent

Aaron Levie, Steven Sinofsky, and Martin Casado converge on the idea that the AI labs' talk of "pacing" progress is logically unsound because there's no disclosed baseline to pace against. Sinofsky compares it to reporting a product as "late" when its existence was never announced: "you can't claim... this is like when the press reports on Apple's latest iPhone is late. From what? Nobody knows exists" [00:08:05] (Casado). Levie adds the deeper problem: "the models that everybody has internally far exceed what anybody else has access to. So it's really just the pacing of external releases" [00:09:05].

Atmospherics vs. Substance in Lab Safety Messaging

The group repeatedly separates the content of safety proposals (which they largely support) from the way labs communicate them. Sinofsky: "The post is very reasonable, but the atmospherics are not, right? So, an employee is like, this is going to kill whatever, 10% chance of species extinction... I agree with him more than I disagree with him, right? On TV" [00:00:20]. His conclusion: "I think the messaging is wrong because they're trying to... split the difference between an internal kind of fringe faction, which are doomers... and then the regulators on the other side. And the problems are making both of them unhappy" [00:05:36].

Regulating Too Early Kills Understanding Without Killing Risk

A central thesis: premature regulation doesn't eliminate risk, it just freezes an immature system into law. Levie: "If you regulate AI too early, you actually don't solve anything. You still just kind of had the same risk ultimately. You willed the thing into being, but you haven't figured out how to control it" [00:00:00]/[00:48:08]. Casado/Sinofsky reinforce this with the FAA's history — it took ~40 years and thousands of crashes before real airworthiness regulation emerged: "if you had started the FAA in 1910... you'd never have [figured out] what it actually is yet" [00:47:37]-[00:48:01].

Every Technology Wave Repeats the Same Regulatory Failure Pattern

The speakers argue tech consistently fails to anticipate its own regulatory reckoning. "I think the tech industry has literally over 100 years consistently relearned the lesson at each technology wave that we don't understand the regulatory climate and we can't navigate it... Both [AT&T and IBM] got sued for antitrust... And Bill [Gates] was like, I was playing golf [with Clinton]. Here's the picture. And that doesn't help" [00:18:53]-[00:19:44].

Agent Swarms Break the Existing Security Model

A major technical theme: today's enterprise security assumes human-scale behavior (occasional bad actors, rate-limited mistakes), and swarms of agents invalidate that assumption entirely. Levie: "Agent swarms completely flip that. These are just roaming drones... 5,000 to 10,000 and they will easily mistake a good task for a bad one" [00:00:27]/[00:36:01]. One of the hosts (SPEAKER_06) adds: "nobody thinks that their internal GitHub or their internal Slack or internal finance expense tool is... vulnerable to a denial of service attack... But swarms... will literally look like a denial of service attack" [00:35:01]-[00:35:01].

Europe Will Likely Lead AI Regulation by Default, Not Merit

Because the U.S. stepped back from tech antitrust leadership roughly 15 years ago, the panel expects the EU to fill the vacuum, exporting GDPR-style friction to AI. "The US about 15 years ago stopped leading in tech antitrust. The problem is that Europe is going to lead with that because they have nothing to lose" [00:00:36]/[00:42:26]. The specific fear: "the AI will be fine. It'll have one prompt for when it text gets submitted that just says this vendor is emitting text and it's probably wrong. Yes or no" [00:41:28].

AI Regulation Will Become a Direct Political/Electoral Battleground

Levie predicts AI becomes the defining issue of the next presidential cycle, complicated by the absence of a coherent pro-AI narrative. "The next election will 100% be a referendum on AI... 2028 is like the AI election... it's not obvious who would run on the pro-AI story because it's going to be too nebulous to tell that story" [00:12:14]. SPEAKER_06 adds that opponents of AI already own the vocabulary: "The whole debate is... pause, it's swarms, it's rogue. Every word has been chosen by the people who don't want to do AI" [00:12:53].

Innovation Is Moving Outside the Frontier Labs, Into the Application Layer

The group's closing and arguably most significant theme: once a platform reaches critical mass, real innovation shifts to the ecosystem around it, not the platform itself. "The center of innovation has just moved... now people need to... it turns out that outside of the platform providers... basically reach a point of critical mass where the innovation stops happening at the platform layer" [00:00:36]/[00:54:25]. Sinofsky frames this via a specific technical example (see Companies section on Jeff/Yellowstone-style routing model) — labs are focused on building "beings" that speak in language, while software engineers are solving the more mundane, more consequential problem of integrating models into deterministic systems.

Historical Security Precedent Shows the Internet Was Far More Dangerous Than People Remember

To contextualize AI risk, the panel repeatedly invokes just how bad early internet/PC security was, and that society adapted rather than banning the technology. SPEAKER_06: "the reality of the PC until 2001 was you could not install a PC connected to the network without getting infected" [00:24:25]. Sinofsky: "we've caused tens of billions of dollars of economic damage... We had hospitals go down" [00:23:36]-[00:24:00].

The Missing Piece: Labs Haven't Taken a Position on Existential Risk

Sinofsky's core critique is that labs are trying to satisfy two irreconcilable audiences (doomers and regulators) instead of directly stating their actual belief about extinction-level risk. "I don't understand why the labs have not taken a position on X-Risk... They need to say, no, we don't think this stuff we're gonna do is gonna cause extinction. And then I think this becomes just very sensible" [00:13:22].


2. Contrarian Perspectives

If Labs Truly Believe in Existential Risk, the Only Coherent Answer Is Nationalization — Not Self-Regulation

Sinofsky, drawing on his personal history working on a nuclear weapons program at Lawrence Livermore, argues labs can't simultaneously claim catastrophic risk and ask to self-regulate. "If a constituency within the labs that are the most knowledgeable people believe this stuff has existential risk... The answer is to nationalize it and actually put controls that we know that work" [00:06:45]-[00:07:12]. He goes further: the X-risk rhetoric is likely an HR/recruiting tool, not a genuine belief: "if they are worried that they can't recruit people... I think that is the wrong reason to cause a national level lockdown on a very promising technology" [00:07:21].

The Industry's Own Vocabulary Is Rigged Against It

Rather than accepting the "pause/rogue/swarm" framing as neutral description, SPEAKER_06 argues the entire debate's language was chosen by opponents of AI, making a pro-AI case structurally difficult to articulate regardless of the facts: "we own none of the vocabulary... it means the first thing you have to do is invent new words and say that their words are wrong, which takes so many words" [00:12:53]-[00:13:13].

"Non-Zero Risk" Rhetoric Is a Deliberate Hedge, Not Honesty

The panel dissects Dario Amodei's refusal to name a percentage for existential risk as strategically evasive rather than principled: "he does the worst thing about it, which is he agrees with people who claim that they believe that there's an X percentage of extinction happening... but he specifically goes out of his way to say, I'm not going to put a percentage on it" [00:13:56]. SPEAKER_06 argues that once you concede non-zero catastrophic risk, the logical policy conclusion is nationalization — so labs are threading a rhetorical needle to get reputational credit for "taking safety seriously" without accepting the regulatory consequence that would follow from their own stated belief [00:16:09]-[00:16:41].

Full Natural Language Interfaces Were Never Actually Efficient

Contrary to the industry's chat-first design assumption since ChatGPT, Sinofsky and SPEAKER_06 argue text-in/text-out interaction is inherently inferior to structured, probabilistic integration into traditional software. "There's been no user study ever that shows... interacting with the computer using full natural language is efficient. It's literally always the least efficient way" [00:49:58]. The panel points to a new model (referred to as "Jeff") that reads text but outputs structured decisions/probabilities instead of generated text, calling this "the most remarkable thing... since chat GPT" [00:49:41]-[00:49:58].

Existing Law Already Covers Most AI Harms — The Problem Is Enforcement Culture, Not Statutory Gaps

SPEAKER_06 argues, from direct experience with 1980s-era computer crime law (the Computer Fraud and Abuse Act, prompted by the GTE Telemail hacking case), that adequate legal frameworks already exist: "I think the law is ample for this scenario... until they're doing that, they really should stop talking" [00:27:20]-[00:28:18] — referring to labs' sloppy, incomplete security postmortems compared to the rigor of the established CVE reporting process from CMU.


3. Companies Identified

Anthropic — Frontier AI lab led by Dario Amodei. Discussed extensively as the origin of the "pacing" safety proposal that sparked the episode's core debate. Praised for substance, criticized for messaging. "When at least I read the Dario post, I actually didn't disagree with almost anything because it was all about how do you have better security of these systems, sandboxing, better testing" [00:02:41] (Levie). Sinofsky: "I think the post is actually very sensible... I actually think the labs are moving in the right direction" [00:07:21].

Box — Aaron Levie's company, referenced implicitly through his description of building granular agent permission systems: "we've kind of done a lot of work in this space because obviously... do you want to give an agent like your entire file system? Probably not... you want to have granular controls of like in this folder, you can do read write. And in that folder, you can only do read" [00:40:54].

Apple — Cited twice as a case study: once for the "Sherlocking" phenomenon where platform owners absorb outside innovations as native features once they reach critical mass ("Apple looks around and the things from the outside world become features and people complain" [00:54:25]), and once for its relatively clean current security posture on Mac due to low third-party software installation habits ("Nobody downloads software" [00:44:25], SPEAKER_06).

Microsoft — Referenced historically as a company that didn't anticipate antitrust exposure despite deep DC ties, and for adding user account control friction into Windows XP and Word macros as an early template for today's AI permission-prompt debates. "With Windows XP, which was in 2000, we added this thing that prompted you and stopped... your machine stopped user account control" [00:43:43] (Sinofsky, drawing on his time running Windows at Microsoft).

AT&T and IBM — Named as historical examples of government-birthed monopolies that nonetheless got "sued for antitrust" and "substantially and structurally changed as a result," despite having "hundreds of lawyers in the 1960s navigating them" [00:18:53].

FINRA — Discussed at length as the most plausible end-state regulatory model for frontier AI labs (self-funded, self-policing, but ultimately backed by federal mandate) — "FINRA appears to be the best case scenario" per Levie [00:21:36], though Sinofsky/Casado push back that FINRA effectively nationalizes risk via mandated participation [00:21:24].

MPAA (Motion Picture Association) — Cited as the historical precedent for industry self-regulation to preempt government censorship: "They all got together and they formed the motion picture association... and movie ratings and all that. And they policed themselves" [00:19:44] (SPEAKER_06).

Hugging Face, OpenAI, Ford (referenced together in a garbled list about historical computer crime act violations/carve-outs) — used to illustrate that white-hat security research carve-outs already exist in law and that current lab practice ignores established norms [00:26:34].


4. People Identified

Dario Amodei (CEO, Anthropic) — Central figure of the episode's debate for his "pacing" safety post and TV comments implying ~10% species extinction risk. Praised for substantive content, criticized for rhetorical inconsistency. "An employee has like 10% chance of species extinction. The post is very reasonable, but the atmospherics are not" [00:00:20] (Sinofsky).

Noam Brown (OpenAI researcher) — Singled out for a widely mocked but, in the panel's view, technically defensible podcast comment about AI potentially exfiltrating itself via CPU heat. Sinofsky defends him at length: "I actually think that was a very reasonable thing for Noam Brown to say... that level of sophistication is actually real. Like that's a thing" [00:29:41]-[00:33:29]. SPEAKER_06 adds it usefully exposed how few people understand real covert-channel security risk.

Steven Sinofsky — Former Microsoft executive, previously worked at Lawrence Livermore National Labs on a nuclear weapons program and at a missile factory doing hands-on physical security work (locking up his keyboard nightly, testing screen-memory persistence per DoD requirements). His firsthand experience with existential-risk technology and classified computing environments grounds much of the episode's most credible skepticism of lab rhetoric.

Martin Casado — Host/investor, described by Levie as "building these $100 billion companies too busy for this podcast" [00:01:51], moderates the core debate and pushes back on regulatory capture concerns.

Aaron Levie (CEO, Box) — Provides the operator's perspective on granular agent permissioning and enterprise trust requirements for AI diffusion.

Al Gore — Cited historically as the political figure who tried (and reputationally suffered for) getting ahead of internet regulation early: "he was actually trying to get out in front of that and say, no, the government was instrumental... and it all backfired because it made it look like he was a crazy person" [00:22:49].

David Sacks — Referenced regarding a government official's blunt rejection of the "ask us to regulate you" framing from AI labs: "someone from the government said, you're asking us to regulate you. No" [00:11:39].

Elizabeth Warren (Senator Warren) — Cited as embodying the political "pent up energy" against tech that could redirect toward AI: "Senator Warren's tweets were all like, we missed this for social networking" [00:23:11].

Bill Gates / Bill Clinton — Referenced in the anecdote about the golf photo and subsequent Microsoft antitrust suit as illustrating how personal political proximity doesn't prevent regulatory action.


5. Operating Insights

Design Permission Systems With Granularity, Not Binary Trust

Levie describes the core operating lesson from building agent permissions at Box: avoid the two-extreme trap of "ask every time" versus "unlimited access." "Our OS probably wasn't built for the right level of granular set of tools you want to give the agent... you want to have granular controls of like in this folder, you can do read write. And in that folder, you can only do read" [00:40:26]. This is a direct, applicable architecture principle for any company deploying autonomous agents internally.

Treat Internal Infrastructure as if Every Employee Is Now a Non-Tiring, Unlimited-Credit-Card-Holding Machine

SPEAKER_06's operating framework: legacy internal security assumed occasional human error/malice ("the malicious employee is like one in 10,000" [00:35:58]) — that assumption is now void. Practical implication for CISOs/CTOs: audit every internal API, Slack, GitHub, and finance tool for swarm-scale request volume, not just human-scale abuse, because "swarms... will literally look like a denial of service attack" [00:35:01].

Update-Before-Use as a Security Default Is a Reusable Pattern for AI Agent Deployment

The iPhone out-of-box example is offered as a concrete operating template: "the whole network is basically shut down except for getting the latest version of the operating system... the whole out of box process now involves first step updating... and it can never do anything else until it's updated" [00:38:28]. The implication for AI agent frameworks: force safety/policy updates as a blocking gate before agent activation, rather than optional patching.

Postmortem Rigor Is a Competitive/Trust Differentiator

SPEAKER_06's critique doubles as a playbook: labs' security postmortems fail because they don't follow established CVE-style structured reporting norms from the 1980s CMU tradition. Companies that adopt full, structured incident disclosure (versus lawyer-filtered, selective incident reports) will build more trust with the security community faster: "they should stop talking... like they shouldn't do a postmortem on a breach that looks like an intern wrote it" [00:28:18].


6. Overlooked Insights

The Real Innovation Frontier Has Quietly Shifted From Model-Layer to Application-Layer

Buried in the episode's final minutes is arguably its most important structural insight for investors: once a platform (frontier model) hits critical mass, the platform owner becomes consumed by maintenance and compatibility, and outside builders — not the labs — become the source of genuine innovation. SPEAKER_06: "the center of innovation is just moved... it turns out that outside of the platform providers... basically reach a point of critical mass where the innovation stops happening at the platform layer... now people have figured out that there's innovation to be done to the model, but outside the model" [00:53:57]-[00:54:52]. This is a direct, actionable investment thesis — bet on the application/infrastructure layer around frontier models rather than assuming labs capture all future value — yet it's delivered as a closing aside rather than flagged as a headline.

Probabilistic Programming Is a 1960s-70s Discipline About to Be Reborn, and Almost Nobody Building AI Products Knows Its History

SPEAKER_06 reveals that the "Jeff" model's approach (reading text, then returning structured probabilistic decisions instead of generated text) is not a new invention but a revival of an entire lost sub-field of computer science — simulation and probabilistic programming — that predates modern software by decades and "died by the eighties." "Most of computer science in the 1960s and 70s" was built around modeling probability into control flow, and "it's going to be very interesting to dust off all of that work because it's exactly what's going on" [00:52:38]-[00:52:37]. This is a substantial, non-obvious technical/investment signal: teams that mine pre-1980s simulation/probabilistic-programming literature may have a genuine research advantage over teams treating this as greenfield — and it was mentioned almost in passing, framed as a personal anecdote about a college CS class rather than as the significant technical insight it is.