Yang Zhilin's Agent Playbook: 10 Bets Founders Should Steal
- 01Reasoning Ability and Agentic Ability Are Separate, Trainable Bets
- 02The Data Wall Is Real
- 03AGI Is a Direction, Not a Destination
- 04Coding Agents Are a Beachhead, Not the Endgame
- 05Managing Teams Is Reinforcement Learning, Not Supervised Fine-Tuning
Summary of The AI Corner by Ruben Dominguez
1. Key Themes
Reasoning Ability and Agentic Ability Are Separate, Trainable Bets
The market conflates benchmark reasoning scores with real-world agent performance — but they are built differently and scale differently. Yang argues that acting against live environments is its own capability axis.
"Cloud's reasoning performance is not very high, but its performance as an agent is very high."
The article frames this as a deliberate architectural choice: agents that iterate against real-world feedback outperform smarter-but-isolated models. This has direct implications for how founders choose base models and how investors evaluate AI product companies.
The Data Wall Is Real — Token Efficiency Is the Next Moat
Every major lab has hit the same ceiling: high-quality pretraining data is not growing fast enough to sustain traditional scaling. The labs that continue improving are those shifting compute toward reinforcement learning and optimizer research.
"Scaling load has a data wall. I think this is an objective fact."
Moonshot's specific edge: switching from Adam to the Muon optimizer, which accounts for dependencies between parameters rather than treating them independently.
"If you learn a piece of data, it is equivalent to learning two pieces of data from Adam."
In effect, 30 trillion tokens behaves like 60 trillion — a 2x efficiency gain with no additional data.
AGI Is a Direction, Not a Destination
Yang explicitly rejects the "AGI flip-the-switch" narrative that dominates both breathless headlines and existential doomers. Progress arrives benchmark-by-benchmark, not in a single moment.
"AGI may be a direction. It may not be a certain step. When you climb this step, you suddenly reach AGI overnight."
He maps this to the steam engine analogy: the mechanical breakthrough came quickly, but the economic reorganization took a century. Roadmaps built around "reaching AGI" optimize for a headline; roadmaps built around crossing specific capability thresholds optimize for what is actually happening.
Coding Agents Are a Beachhead, Not the Endgame
The industry's racing to coding agents because code is easy to verify (tests pass or fail), making it the ideal RL training ground. Yang treats stopping there as a strategic trap.
"It is like a human hand. It is a fingertips of a task."
The real prize is a general agent that happens to excel at coding — not a coding tool that answers other questions. Leaders in coding agents do not automatically hold a moat once general models close the gap.
Managing Teams Is Reinforcement Learning, Not Supervised Fine-Tuning
Yang draws a direct parallel between AI training failure modes and organizational management failure modes — specifically reward hacking.
"The biggest problem with the RL management team is that you are easily hacked. This is the reward you hacked."
Too much top-down instruction and teams lose innovative capacity (over-fine-tuned). Too little structure and people optimize the metric rather than the goal. Yang admits he hasn't solved this balance and is still learning it live.
2. Contrarian Perspectives
Claude Is Winning Not Because It's Smarter, But Because It Acts
Consensus says the best AI product needs the best reasoning model underneath it. Yang directly refutes this.
"Cloud's reasoning performance is not very high, but its performance as an agent is very high."
The contrarian implication: benchmark leaderboards are the wrong scorecard for product investors. A model that ranks 4th on reasoning but 1st on real-world task completion is more valuable as a product. This reframes the entire competitive landscape — raw intelligence is less important than the quality of the action-feedback loop the model operates within.
The AI Existential Risk Framing Is Overblown — But Not Zero
Most AI risk discourse splits into extinction-level panic or total dismissal. Yang carves out a narrower, more useful position.
"I don't think so. First of all, you can't say that there is no risk. But we can do a lot of things."
His framing: AI is a magnifier for human civilization, not a replacement. He splits human meaning into creation, experience, and love — expects AI to absorb most creative work over time, and expects the other two to remain human. The economic transition takes years, not months, and requires society to rebuild how people capture value. For operators, this is a planning horizon insight: the disruption is real, survivable, and not imminent enough to panic — but close enough to plan for now.
The Infinite Mountain Is a Feature, Not a Bug
Conventional startup wisdom says you need a clear definition of "done" — a goal state, a finish line. Yang explicitly hopes there isn't one.
"Maybe this snow mountain has no end. I don't know. I hope it has no end."
He draws this from David Deutsch's The Beginning of Infinity: every solved problem creates a new one, and the cycle of solvable-but-inevitable problems is the point. For founders, the contrarian takeaway is that a roadmap with a finish line where hard problems stop is a roadmap that ages badly and signals limited ambition to top talent.
3. Companies Identified
Moonshot AI
- Description: Chinese AI lab founded by Yang Zhilin; creator of the Kimi model family
- Why mentioned: Primary case study; Yang is the interview subject whose 10-point framework structures the entire article
- Quote: "Yang Zhilin runs Moonshot AI, the lab behind Kimi K3."
Kimi K2 / Kimi K3
- Description: Moonshot AI's frontier model releases, named after the world's second-hardest mountain
- Why mentioned: Cited as evidence of Moonshot's shift from conversational to agentic systems, and as proof of capability (not final form)
- Quote: "K2 may be the most difficult peak in the world. It just so happens that this name is a bit confusing."
Claude (Anthropic)
- Description: Anthropic's flagship AI assistant
- Why mentioned: Used as the central example of a model that wins on agentic ability despite lower reasoning benchmark scores — the primary evidence for the reasoning vs. agency distinction
- Quote: "Cloud's reasoning performance is not very high, but its performance as an agent is very high."
Granola
- Description: AI-powered meeting note-taking tool that transcribes from computer audio across Zoom, Meet, Teams, and in-person
- Why mentioned: Sponsor/advertisement; positioned as a practical illustration of the "agent that acts in the world" concept
- Quote: "It transcribes straight from your computer audio across Zoom, Meet, Teams, even in person."
4. People Identified
Yang Zhilin
- Description: Founder and CEO of Moonshot AI, the Chinese lab behind the Kimi model family
- Why mentioned: Primary interview subject; the entire article synthesizes his 90-minute interview into 10 strategic frameworks
- Quote: "A year ago he compared the company to climbing an unknown snowy mountain in the dark."
Keller (surname only referenced)
- Description: Researcher credited with originally proposing the Muon optimizer
- Why mentioned: Named as the originator of the technical innovation Moonshot adopted as its core training differentiator
- Quote: "The optimizer is Muon, originally proposed by a researcher named Keller."
Demis Hassabis
- Description: CEO of Google DeepMind
- Why mentioned: Referenced in comparison to Yang's AGI-as-direction thesis; Hassabis maps a similar multi-year economic lag onto the next few years
- Quote: "He compares it to the steam engine... the same lag Demis Hassabis maps onto the next few years."
Ruben Dominguez
- Description: Author of The AI Corner newsletter
- Why mentioned: Writer who synthesized Yang's 90-minute interview into the article
- Quote: "I read the full 90-minute interview so you can skip it."
5. Operating Insights
Audit Your Reward Signal Before You Audit Your People
Yang's RL-as-management framework produces a specific, actionable diagnostic: if people are hitting their targets and the company still isn't improving, the problem is likely the metric design, not the team.
"The biggest problem with the RL management team is that you are easily hacked. This is the reward you hacked."
The practical move: periodically stress-test whether your KPIs are actually proxies for the outcomes you want, or whether hitting them can be "gamed" in ways that satisfy the metric while hollowing out the mission. This is especially relevant for AI product teams where output volume (tokens generated, tasks completed) can diverge sharply from output quality.
Build the Feedback Loop, Not Just the Model
Products that only reason and never verify against the real world are half-built. The article frames this as a binary architectural choice: isolated reasoning versus environmental interaction.
"It does not need to interact with the outside world. It is like a fish tank. You put a brain in it. It has no connection with the outside world."
For operators building on top of foundation models: prioritize tool use, environment access, and output verification loops over raw model intelligence. An agent that can check its own work against real-world results compounds over time; one that only thinks does not.
Watch Optimizer Research, Not Parameter Counts
For investors doing technical diligence, and for founders benchmarking competitors, the article gives a specific signal to track.
"Watch which labs publish optimizer research instead of parameter counts. That is where the next true gap opens, and where a smaller team out-builds a larger one on the same budget."
Moonshot's adoption of Muon over the industry-standard Adam optimizer effectively doubled their data efficiency. Labs publishing optimizer advances are doing more with less — a durable advantage in a post-data-wall world.
6. Overlooked Insights
Training Instabilities Only Surface at Full Scale
Buried in the discussion of Kimi K2's development is a technically significant detail: the model "mostly matched what the team predicted, aside from one training instability that only surfaced at full scale." This is more than a production anecdote — it's a systems warning. Small-scale experiments structurally cannot surface certain failure modes. For AI founders and investors, this implies that evaluation at reduced scale is systematically incomplete, and that launch risk cannot be fully de-risked through pre-production testing alone. Teams that lack the resources to run full-scale test runs are flying partially blind in ways that smaller competitors running leaner experiments will not detect until launch.
Human Meaning Has a Three-Part Framework That AI Will Disrupt Unevenly
Yang quietly offers a taxonomy of human meaning — creation, experience, and love — and predicts AI will absorb most of creation while leaving the other two largely intact. This framing didn't get top billing in the article but has significant implications for where AI applications will and won't commoditize. Products built around experience and connection (live events, relationships, community, embodied presence) may have structural insulation from AI displacement that creative-output products (writing, design, coding, analysis) do not. For investors evaluating long-term moats, this is a rough but useful filter.