Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/LENNY'S/Anthropic’s first technical PM o…
POD
// EPISODE
LENNY'S

Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne Penn

DATE July 26, 2026SOURCE LENNY'SPARTICIPANTS DIANE PENN, LENNY RACHITSKY
// KEY TAKEAWAYS6 ITEMS
  1. 01Coding Was the Hidden Inflection Point That Made Anthropic
  2. 02Frontier Products and Frontier Models Are Mutually Dependent
  3. 03Evals Are the New PRDs
  4. 04Emerging Capabilities Are Discontinuous and Unpredictable
  5. 05The Labs Model: Strong Conviction on Theme, Loose on Prototype
  6. 06Claude's Non-Compliance Is a Feature, Not a Bug

1. Key Themes

Coding Was the Hidden Inflection Point That Made Anthropic

Before anyone drew the connection, Diane identified that users were using Claude not just for autocomplete but for long-form code generation — and pushed to train Opus 3 on it. This turned out to be the competitive differentiator that brought developers to Claude early.

"In 2023, when I started, nobody said Anthropic and Claude and coding in the same sentence... I saw people were starting to use these models not just for code autocomplete, but actually writing long form code. And it's that an opportunity for us to train Opus 3 to be better at. It ended up being a relatively smaller change from a training perspective, but it ended up helping us differentiate in the early days competitively." 00:00:00

Frontier Products and Frontier Models Are Mutually Dependent

Diane articulates a rarely stated but critical insight: a great model without a great product surface cannot achieve adoption, and a great product without a frontier model cannot demonstrate magic. The Opus 4.5 / Claude Code flywheel is the clearest proof case.

"You need frontier products in order to have frontier models and for people to feel the magic of frontier models... Opus 4.5 wouldn't have had that moment without a product like Claude Code. And Claude Code wouldn't have had that type of adoption accelerated without Opus 4.5." 00:12:38

Evals Are the New PRDs — The Product Workflow Has Fundamentally Changed

The core artifact of the product development process has shifted. Rather than writing specs that describe what to build, research PMs now write evals — structured test sets that define what "good" looks like — as the primary mechanism for driving model improvement.

"For my team, the way to drive user value is to figure out the right user feedback, the evals. We actually have a saying on the team of evals are the new PRDs." 00:00:52

"The way to even access a user pain point is different, right? In the past, we might do a user interview... here, you have to sweat the tokens as much as you sweat the pixels." 00:42:31

Emerging Capabilities Are Discontinuous and Unpredictable — Which Has Safety Implications

Diane points to the original scaling law papers to make a non-obvious point: while loss decreases smoothly with scale, new capabilities appear as sudden jumps — not gradual improvements. This is both a product opportunity and a safety challenge.

"As you add in more data and you train the models with more compute, you essentially see these actually discontinuous emerging capabilities jump... Unless you have the evals, unless you have the systems to test, these jumps might actually happen and you don't know." 00:18:17

The Labs Model: Strong Conviction on Theme, Loose on Prototype

Anthropic's internal labs team operates on a specific operating principle — have a firmly held thesis about the space, but iterate rapidly on execution. Bets that don't work immediately may be revisited in one to two model generations.

"You can be very strongly held opinion about the theme or the area, and then more weakly held about the exact prototype... We have a thesis and it might not work yet. And so we then might revisit it in one to two model generations." 00:24:28

Claude's Non-Compliance Is a Feature, Not a Bug

The alignment and safety work — giving Claude a constitution and the ability to push back — is counterintuitively what makes Claude more useful and more interesting than competing models. A compliant AI is a less useful AI.

"What you don't want is an AI that just agrees with you. What you want is this technology to actually augment and grow and get to a better outcome. And so sometimes it's having Claude push back that makes me better. Like a co-worker, I want somebody to push back when my ideas are not fully formed." 01:03:21

"A thinking partner doesn't just agree with you. It should add to you. And you should come away at the end of the day having better ideas because you worked with Claude. That should be the hero goal, not just making your ideas 10% better." 01:06:26

Forward Compatibility as a Product Design Constraint

Diane's team explicitly asks: what happens when Claude 8 exists? Then builds today so that the product is forward compatible with that future. This prevents building experiences that model improvements will obsolete.

"One thing I ask the team frequently is, let's say Claude 8 comes around. What changes in what users do? And then what does that mean for how you're building today? Is it going to be forward compatible to that experience?" 00:34:29

The "Product Overhang" Opportunity in AI

Diane uses the term "product overhang" to describe the gap between what current models can do and what products have actually been built on top of them. The implication: the current models are more capable than what users are experiencing.

"There's like product overhang and user overhang, like to maybe put it in our PM language, even on today's models. And I think there's like a lot that we could be exploring on our current opuses and definitely with Fable, for example." 00:19:27

AI Writing Quality Is Lagging — But Deliberately

The model's writing quality has been a lower priority because the team was focused on more pressing capability jumps like agentic behavior and tool use. Now that those are more mature, writing quality is actively being addressed as the next rough edge.

"The technology is jagged edged... sometimes when the models were good at writing, but not agentic, our thesis is how do we make the models more agentic or call the right tools? Now that that's improved a bit, then it's, well, now these other areas actually become more of the rough edges." 01:09:20

Experimentation Is a Team Sport, Not an Individual Activity

One of the most underappreciated insights in the episode: the communal discovery model — where people share what they're trying and iterate together — is faster and more effective than individuals trying to figure out AI alone.

"The way you would see magically is different users or different folks on the team coming up with an idea and then other people trying different variations of that idea. And then within maybe 10 or so requests, there was something magical or potentially in a use case that emerges." 00:22:14


2. Contrarian Perspectives

Being the Underdog Against OpenAI Was Not a Disadvantage — It Was a Filter for Mission Alignment

When everyone assumed Anthropic had no chance, only people who deeply believed in the mission joined. This created an unusually high-conviction, culture-aligned founding team that became a durable competitive advantage.

"A big portion of it was the culture was really strong... really do walk the walk of the mission and the culture and the values. And the energy was very much like a startup." 00:03:44

Token Spending Is the Wrong Frame — Experimentation Rate Is What Actually Matters

Gary Tan's "spend $100K a year in tokens to live in 2028" frame is popular, but Diane pushes back on it. Token spend is the input; the output is experimentation frequency. Optimizing for the metric rather than the outcome is a mistake.

"I take more of like an almost product lens. It's almost like token spend is more the input. And really the output is what you described of experimentation... I feel like that might be the better framing of the outcomes. And therefore there might be different ways of achieving that outcome." 00:20:37

Frontier Model Restrictions Create an Unexpected Moat for AI Labs

As models become powerful enough to require safety controls and restricted access, the labs building those models get privileged access to the most capable tools before anyone outside can use them. This creates a compounding internal advantage that was not anticipated.

"It creates this unfair advantage within the labs to have access to the best stuff that other people can't get outside of your control... It's a really interesting, this new feedback loop that's going to start where models that are so advanced are only accessible to certain companies." 00:37:28 (Lenny Rachitsky)

AI Writing Being "Obviously AI" Is Not Always a Problem to Solve

Rather than striving for AI writing that is indistinguishable from human writing, the more important question is who is verifying the output, not who wrote it. For many use cases, AI-authored text that is clearly AI is preferable.

"Maybe the lens is more around like verifiability or who's verifying the output. Like who's signing off. Maybe less around who's writing, but who's verifying, who's signing off. That becomes like more what matters than who's writing it." 01:08:29

PMs Are More Necessary Now, Not Less — But for Completely Different Reasons

Conventional wisdom says engineering-empowered AI reduces the need for product managers. Diane's view is the opposite: as building becomes trivial, the hard problem of knowing what to build and deeply understanding users becomes more valuable, not less.

"Do we still need PMs when the models are so capable, when engineers are leaning in? And I think the role of people who are user-centric, who go into the details of understanding what users are trying to accomplish, bubbling that up in an actionable manner and doing the relentless work to do that — I actually think we need more of that." 01:22:44


3. Companies Identified

Anthropic AI safety company and developer of Claude. The primary subject of the episode — discussed at length for its culture-first approach, labs model, research-to-product integration, and its trajectory from underdog to reportedly $50B ARR.

"We did four models in the whole year or four series of models [in 2024]. And I think we did more than that volume in just Q2 of this year." 01:17:01

WorkOS Enterprise authentication and compliance infrastructure provider. Described as powering OpenAI, Anthropic, Cursor, Vercel, Replit, Sierra, Clay, and described as "Stripe for enterprise features."

"Literally every startup that I'm an investor in that starts to expand upmarket ends up working with WorkOS." 00:08:10 (Lenny Rachitsky)

Mercury Fintech banking platform for entrepreneurs. Mentioned for having an MCP server, a CLI, an API, and a new AI-powered financial operator interface called Command.

"Does your bank have an API, a terminal native CLI, or an AI-ready MCP server? I don't think so." 00:38:43 (Lenny Rachitsky)

Y Combinator Startup accelerator, referenced through Gary Tan's insight about token spending as a proxy for living in 2028.

"If you're willing to spend $100,000 a year right now in tokens, you are living the way somebody in 2028 is going to live." 00:01:02 (Lenny Rachitsky, citing Gary Tan)

JP Morgan Chase Diane's early employer as a high-yield bond trader; mentioned for the lessons it taught her about conviction, authentic self-expression, and meritocracy of ideas in a male-dominated environment.

"Even if I was the most junior person, even if I may look different — the best ideas and having conviction in the best ideas, regardless of all of those other factors — is the most important thing." 01:30:20

Amazon Mentioned as Diane's employer prior to Anthropic, establishing her six years in AI overall.

"This is year six of me working in AI. So Amazon and then Anthropic." 01:19:26


4. People Identified

Diane Penn Head of product for AI research and labs teams at Anthropic; first technical PM at Anthropic. Joined when the product team was five engineers, has shipped every model from Claude 2 through Fable, and incubated Claude Code, MCP, Claude Design, computer use, tool use, and reasoning.

"She's helped ship every model at Anthropic from Claude 2 through Fable. She's also helped incubate and launch Claude Code, MCP, skills, Claude design, and also core capabilities like computer use, tool use, and reasoning." 00:01:35 (Lenny Rachitsky)

Ben Mann Co-founder at Anthropic and labs team leader. Praised for setting an extraordinary innovation vision and pushing teams to think 10x/100x on bets. Also noted for championing Montessori education and the memorable line that "this is the most normal it's ever going to be."

"Ben sets an incredible vision and pushes people to think about the 10x, 100x of the idea." 00:26:12

Dario Amodei CEO of Anthropic. Highlighted for his prescient predictions about AI-driven software engineering at a time when most dismissed the idea as impossible.

"Dario — we can transform software engineering. And I think going in that direction, you learn so much." 00:33:19

Gary Tan President of Y Combinator. Cited for the "live in 2028 now" token-spending thesis.

"Something Gary Tan's been talking about... If you're willing to spend $100,000 a year right now on tokens, you are living the way somebody in 2028 is going to live." 00:01:02 (Lenny Rachitsky)

Mike Krieger Co-founder of Instagram; now on the Anthropic labs team. Mentioned as a prior podcast guest and example of the caliber of talent in labs.

"We've had Ben Mann on the podcast, Mike Krieger, whom both work on labs now." 00:23:33 (Lenny Rachitsky)

Eric Ries Author of The Lean Startup and the recent book Incorruptible. Recommended by Diane for his thinking on sustaining company values through culture metrics rather than revenue metrics alone.

"I loved some of the examples around having metrics around culture. If you only measure revenue and then that's kind of how you're goaling against. But if you have other better metrics, that's actually the way to sustain the values you care about." 01:25:27

Fiona Fung Recent Lenny's podcast guest who suggested the burnout question and was credited for insights on how software engineering is becoming lonelier as teams shrink and people work with agents instead of humans.

Referenced by Lenny: "Fiona Fung actually suggested I ask you, who's recently on the podcast... she pointed out it's a lot lonelier now because now we're working with agents instead of other humans." 01:16:11 / 01:20:12

Andrew (Head of Codex App, OpenAI) Referenced as having aligned with Diane that the PRD is not dead, still useful for specific projects. Not fully named.

"I just had Andrew, he's the head of the Codex app at OpenAI... you guys are aligned. PRD is not dead." 00:49:39 (Lenny Rachitsky)


5. Operating Insights

"Think First, Then Spar" — Protecting Cognitive Ownership While Using AI

Diane's personal protocol for avoiding AI-driven homogenization of thought: form your own point of view first, then use Claude as an adversarial thinking partner. Reserve full delegation only for low-judgment, standardized outputs like business reviews.

"What I want to make sure is Claude doesn't take over all of my thinking for me. I might come up with my own POV first and then work with Claude through that... For something like a monthly business review, I actually want it to be standard. I want it to be much more crystallized information in the right way... I want to get to a place where the writing of that is potentially asymmetrically less valuable than the thinking." 01:01:37

Building a Crucial Conversations Skill to Become a Better Manager

Diane built a Claude skill grounded in the book Crucial Conversations and uses it to prep for difficult managerial conversations — treating Claude as a real-time personalized coach before high-stakes interactions.

"I love that book. And so I actually have a skill that helps me figure out, am I going in the right level of detail given the situation at hand and actually helping me be a better manager and better supporter for the team?" 00:59:21

Managers Must Stay Hands-On and Ship — Same Onboarding as Junior PMs

At Anthropic, seniority does not exempt product leaders from hands-on work. Senior PMs and managers go through identical onboarding to junior staff, and Diane explicitly carves out time to personally own one to two workstreams per model cycle to maintain calibration.

"Even for folks that I hire who have more tenured PM experience, the onboarding plans are exactly the same as somebody who is more early career... I always try to carve out a portion of time to actually own one to two work streams when we have models in order to keep my theory of mind, keep my sense of how the models are moving." 00:50:52

Culture Metrics as an Operating Tool — Not Just Values Posters

Directly inspired by Eric Ries' Incorruptible, Diane is actively working on codifying team norms and culture metrics at the team level — measuring culture the same way you'd measure revenue — to make values durable as Anthropic scales.

"If you have other better metrics, that's actually the way to sustain the values you care about. I've been kind of trying to think about how to actually bring that to the team level of like, how do we better articulate, write our norms?" 01:25:51

Hire for Low-Ego Team Orientation, Not Org-Building Ambition

Diane's explicit hiring filter for sustainability in a fast-moving environment: is this person building their own empire or contributing to the team? Low-ego team players are what allow a team to genuinely cover for each other and prevent individual burnout.

"Is this person going to care about their own ego and building out a big org, or are they going to care about contributing to Anthropic and contributing to the impact of the team? Orienting towards folks who are low ego, team oriented — that's a big part of the sustainability." 01:20:51


6. Overlooked Insights

Fallback UX Architecture Is Now a Core Competitive Infrastructure Layer

Briefly mentioned in the context of Fable's safety restrictions, Diane reveals that Anthropic built "fallback UX systems" that automatically route users to Opus 4.8 when Fable is restricted — meaning Anthropic is now engineering graceful degradation paths between model tiers as a product capability. This is not just a safety measure; it is a product architecture pattern that allows frontier model restrictions to exist without destroying user experience. Every AI product company will eventually need this infrastructure, and almost none have it yet.

"Before Fable models, we didn't have a strong, let's say, fallback UXs and systems... We ended up building fallback systems so that users will still get a great response from Opus 4.8 immediately... You'll see us innovating, improving on what we call now the model safeguards package more and more in the coming weeks and months." 00:36:31

Golden Gate Claude as a Proof-of-Concept for Interpretability-as-Product

In passing, Diane describes a 24-hour experiment called "Golden Gate Claude" where Anthropic took a specific interpretable feature from their mechanistic interpretability research (the Golden Gate Bridge feature) and surfaced it directly as a user experience. This was dismissed as reaching only 2,000 people — but it represents a genuinely novel product category: turning interpretability research directly into user-facing product experiences. As interpretability matures at Anthropic and elsewhere, the ability to let users interact with or even dial up specific internal model features could become an entirely new product surface that no one is actively building toward yet.

"We had just published one of our early interpretability research in early 2024... one that really came up frequently that resonated was the Golden Gate Bridge... when you actually essentially dialed up that feature, Claude would obsess about the Golden Gate Bridge... The entire experience we spun up on our Claude.ai website within 24 hours. That to me was one of those hidden inflection points of we were starting to find our identity." 00:05:07