Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/晚点聊 LATETALK/178: 与田渊栋聊 RSI:模型自进化如何到来?
POD
// EPISODE
晚点聊 LATETALK

178: 与田渊栋聊 RSI:模型自进化如何到来?

DATE August 7, 2026SOURCE 晚点聊 LATETALKPARTICIPANTS MANCHI, 晚点团队
// KEY TAKEAWAYS6 ITEMS
  1. 01The Timing of RSI's Emergence Is Tied to Coding Agent Maturity
  2. 02Hands-On Practitioners Are the First to See Paradigm Shifts
  3. 03The Scaling Law Debate Has Direct Implications for Startup Viability
  4. 04RSI Is Qualitatively Larger Than Coding Agents
  5. 05The Key Missing Capability for RSI Is Scientific Creativity, Not More Reasoning
  6. 06RSI Is a Staircase, Not a Binary Threshold

1. Key Themes

The Timing of RSI's Emergence Is Tied to Coding Agent Maturity

RSI as a concept has existed for years, but only became actionable recently because coding agents and base model capabilities crossed a threshold. The key unlock was models becoming strong enough to act as research collaborators rather than just tools.

"I think the main reason is because Coding Agents have become stronger. Also, the model's own capabilities have improved. Once the model's capabilities improve to a certain level, the model's understanding of AI itself achieves some strong breakthroughs — you can then use the model as a researcher. Let it find the weaknesses of the current model, then do better optimization. That's how the closed loop is created." — Tian Yuandong [00:06:10.800]

Hands-On Practitioners Are the First to See Paradigm Shifts

Those who could identify the RSI opportunity earliest were senior researchers who stayed close to the code — not executives thinking at a high level, and not junior engineers executing without judgment. The combination of domain expertise and hands-on practice creates a unique signal-detection advantage.

"If you don't truly treat the model as an intern, or treat it as just a chat object, you won't feel the change. So you have to wait until you really treat the model as a collaborator, work with it on a paper... then you'll realize the model's capability is already much higher than I imagined." — Tian Yuandong [00:09:06.220]

"I think the balance will shift back toward researchers with more experience and seniority... execution used to require people. Now the execution burden has fallen to AI. People now focus more on judgment, logical guidance, direction-setting, and research taste — these things matter." — Tian Yuandong [00:15:28.360]

The Scaling Law Debate Has Direct Implications for Startup Viability

Whether capability curves are smooth and continuous (pure Scaling Law) or step-function/discontinuous determines whether startups can compete with frontier labs. Tian believes the curve has plateaus, which is precisely the structural opening for smaller, research-first companies.

"I think there are plateaus. You discover something new, have a new idea, and the model built from that idea becomes stronger. Then with that strength you get even better ideas. It's a staged, leap-by-leap process." — Tian Yuandong [00:40:12.400]

"If you completely believe in Scaling Law, then startups have no chance — because OpenAI and Anthropic are already so far ahead, and if they accelerate into RSI, they'll always be faster. But I don't think it's that way. Scaling Law may not be everything." — Tian Yuandong [00:41:10.620]

RSI Is Qualitatively Larger Than Coding Agents — Not Just an Extension

There is a tempting but flawed narrative that RSI is simply the next step after Coding Agents on the same rails. Tian pushes back firmly: RSI requires capabilities — research taste, abstraction, creativity — that Coding Agents don't have and that frontier labs don't automatically possess just because they lead on coding.

"I think RSI is much larger than Coding Agents. Its difficulty, and its possible paths, are much greater than Coding Agents today. So just having that pivot doesn't mean you can necessarily do RSI well." — Tian Yuandong [00:27:35.700]

The Key Missing Capability for RSI Is Scientific Creativity, Not More Reasoning

The hardest thing RSI needs isn't reasoning — it's something closer to grokking: the ability to extract deep patterns from sparse data, synthesize new principles from messy phenomena, and generate genuinely novel hypotheses. Current models have this capacity only in domains with abundant existing data.

"This is of course reasoning ability, but it's hard to describe. It's bigger than reasoning... In terms of human ability, most people don't have it — only some very strong researchers do. It's creativity. The ability to extract concepts from complex, tangled surface phenomena, summarize and think about them, and apply them to new problems." — Tian Yuandong [00:31:00.340]

"How AI learns to recognize the world, discover new patterns, generate emergence — since 2013-2014, many theorists tried to explain this: renormalization groups from physics, spin glass models, Neural Tangent Kernels from math, Bayesian methods, Gaussian processes. Many angles were tried, none with satisfying results." — Tian Yuandong [00:34:19.800]

RSI Is a Staircase, Not a Binary Threshold

Unlike autonomous vehicles (where it's 0 or 100), RSI has meaningful intermediate value at every step. Each tier of capability is commercially and scientifically useful in itself, giving RSI companies sustainable intermediate milestones rather than an all-or-nothing bet.

"For RSI, it's not like autonomous vehicles where it's either 0 or 100 — it has many steps. Because even at the first stage of RSI, it can already do many things. It has real practical value already. The second step will unlock deeper things, even bigger. The ultimate goal is AI with a brain like Einstein or Newton, able to gain deep insights from very few samples." — Tian Yuandong [00:36:17.260]

Interpretability Is the Underrated Foundation for Both Capability and Safety

Tian identifies interpretability (可解释性) not as a side concern but as the core research direction that will simultaneously accelerate model capability and make models safer — and he believes strongly it is achievable, against a prevailing sentiment of skepticism.

"I'm probably in the minority, because most people have already given up hope — they think the model is so complex there's no chance of understanding what's going on inside. But I don't think that's right. I believe there must be a good principle that makes the model work. If you truly find that principle, on one hand it makes the model stronger; on the other hand it makes the model safer." — Tian Yuandong [00:59:33.060]

Large Organizational Structure Is Actively Harmful to Frontier AI Progress

Tian offers a structural critique: beyond roughly 150 people, the two-tier (director/executor) relationship that large organizations create slows AI research because it mirrors the slow professor-student feedback loop that AI can now compress to minutes. Flat, small, hands-on teams have a structural advantage.

"The larger and more bloated the organization becomes, the slower AI progress gets. Because AI very much needs people to be hands-on — to truly use it, truly feel it, truly find its problems... Once you have a two-tier structure of thinkers and executors, the speed slows down, just like what I said about supervisors and students." — Tian Yuandong [01:21:07.000]

"Llama 4 and XDR both started falling apart after exceeding 150 people. Beyond 150 people, it breaks down. With a private team of 30, everyone knows everyone, relationships are very simple." — Tian Yuandong [01:22:05.800]

The "Winner Takes All via RSI Feedback Loop" Narrative Has a Structural Flaw

The popular Silicon Valley thesis — that whoever enters RSI first creates an insurmountable compounding advantage — assumes a smooth, continuous curve. Tian argues breakthroughs are discrete jumps; before each jump, everyone is roughly equal, and after the jump, one or two actors leapfrog. This is the structural argument for why startups remain viable.

"Maybe before reaching a certain model level, everyone is roughly equal, doing reasonably well — but then one or two people break through. Once they break through, they reach the next tier, and that tier's AI capability far exceeds the previous tier. The whole level jumps faster." — Tian Yuandong [00:39:12.440]


2. Contrarian Perspectives

The Scaling Law "Smooth Acceleration" Thesis Is Probably Wrong — and That's Good News for Startups

Conventional wisdom in 2025 says whoever has the most compute will compound the fastest into RSI. Tian argues this is incomplete: Scaling Law requires 10x resources for linear gains, and physical constraints (energy, compute) already create hard ceilings. The curve must be stepped, not smooth.

"Scaling Law may not be everything. It may be correct, only that you need 10x resources, 10x data, 10x everything, and ultimately achieve linear growth. But 10x resources already face bottlenecks now. You can't have the entire planet's electricity supply feeding one entity's self-improvement — there will definitely be constraints." — Tian Yuandong [00:41:10.620]

The "Youth Advantage" Narrative in AI Is Reversing with RSI

The dominant narrative since 2022 has been that young, fast-moving people without established habits have a structural advantage in AI. Tian argues RSI inverts this: you need deep domain intuition to know which problems are hard, which signals matter, and to have the taste to direct autonomous AI effectively.

"Before Coding Agents became widespread, young people had an advantage mainly because of their drive and fast execution speed. Many times you didn't need to think too much — you just started doing it. That was basically the past three years." — Tian Yuandong [00:13:32.040]

"RSI's arrival may tip the balance back toward more experienced, senior researchers... The execution burden is now on AI. What matters is judgment, logical guidance, direction control, and research taste." — Tian Yuandong [00:15:28.360]

AI Self-Evolution Does Not Require or Imply Consciousness — and Safety Arguments Based on That Premise Are Flawed

The common fear that RSI-capable AI will develop dangerous self-awareness conflates intelligence with volition. Tian distinguishes these sharply, noting that genius and self-interested desire for power are entirely separable traits — in humans as well as AI.

"I don't think they're the same thing at all. You can think about it — many genius mathematicians and physicists are extraordinary in their domains, but at the same time they have no interest in worldly power. These two capacities — capability and desire — are completely separable. Our training signal is such that AI is trained to serve humans well; under that evolutionary pressure, AI has no way to evolve self-consciousness or self-awareness the way humans have." — Tian Yuandong [00:58:36.700]

Anthropic's "When AI Builds Itself" Work Is Still Just Efficiency Optimization, Not True RSI

Tian gently deflates the hype around Anthropic's June 2025 RSI blog post, characterizing it as scaling up conventional algorithms rather than recursive self-improvement in a meaningful sense.

"Basically it's still efficiency gains — finding various methods to improve Weak-to-Strong performance. The algorithms themselves are fairly standard. It puts many conventional ideas together, and the final score goes up. The reason you need scores is because without them you have no other way to measure current model capability." — Tian Yuandong [00:21:13.960]

The Next "Newton" Could Be AI, Not a Human — and That Is the Entire Point of RSI

This is a minority position even among RSI researchers. Most assume humans will remain the source of scientific breakthroughs and AI will accelerate them. Tian's vision is more radical: the next synthesis of physical law or fundamental principle may be generated by an AI system operating on far less data than current models require.

"The next Newton to appear — it may not necessarily be a human, it could be AI, it could be human-plus-AI. That's exactly why we do AI research. If you want to do this, you must use all the world's computational resources and the strongest minds — both human and AI — to make it happen." — Tian Yuandong [01:05:26.500]


3. Companies Identified

Recursive Super Intelligence (RSI) A startup co-founded by Tian Yuandong focused on recursive self-improvement — building AI systems that can discover new architectures, training algorithms, and optimizations autonomously. Raised $650M Series A at a $4.65B valuation after only 4 months of stealth R&D. Released first results in June 2025: achieved SOTA on NanoGPT speedrun, NanoGPT 5-minute training, and NVIDIA's SOL Exact Bench for kernel optimization.

"This company's purpose is to find a method for AI to self-evolve and self-iterate. To find new paradigms, new models, new training logic, and thereby obtain better AI." — Tian Yuandong [00:04:18.920]

Anthropic Frontier AI lab. Published "When AI Builds Itself" in June 2025 documenting their RSI progress. Tian assessed their current RSI work as still at the efficiency-improvement stage rather than genuine recursive self-improvement.

"Anthropic's article basically still reflects what we discussed — AI enabling AI to self-evolve and self-iterate, getting better and better. The algorithms inside are fairly standard, putting many conventional ideas together." — Tian Yuandong [00:20:17.700]

OpenAI Frontier AI lab. Publicly committed to producing an "AI research intern" by September and fully automated AI researchers by March 2028.

"OpenAI has again made its timeline clear — saying they want to produce AI research interns this year in September, and before March 2028 to achieve truly automated AI researchers." — Host Manchi [00:03:15.940]

Miraendar (Mirandale) New RSI startup, officially announced June 25 (just before this recording). Raised $200M seed round. Team reportedly from DeepMind and Anthropic.

"There's one that just officially announced on June 25th, called Mirandale, and they raised $200M in a seed round. These teams are reportedly from DeepMind and Anthropic." — Host Manchi [00:23:11.780]

Sakana AI Named as another company that has formed a new RSI-focused team as of June 2025.

"The previously established Sakana AI also just formed a new RSI team in June." — Host Manchi [00:23:39.540]

NVIDIA Created SOL Exact Bench — a recently launched benchmark for kernel/operator optimization — which RSI's system achieved SOTA on. Tian notes NVIDIA's interest is straightforwardly that faster AI kernels benefit their hardware.

"For NVIDIA, doing this benchmark is mostly to see — because AI kernels getting faster is good for NVIDIA. So it doesn't matter to them whether it's a human or a system doing it. Better kernels for its hardware means faster speed and higher efficiency — that's always good for them." — Tian Yuandong [00:48:01.740]

AAI (Israeli company) Competitor on the SOL Exact Bench. Led by Sasha, former CEO of Mobileye. A professional semiconductor company with GPU experts. RSI's non-specialist system beat them by 10%.

"The second place had a company called AAI, an Israeli team. Their CEO is Sasha, the former CEO of Mobileye — a very well-established company doing very strong chip-level processing. They had previously supplied chips to Tesla. It's a professional semiconductor company. Their team are professional GPU engineers. They submitted a version, but our results were 10% better." — Tian Yuandong [00:51:22.920]

Mobileye Described as a longtime semiconductor/autonomous driving chip company. Mentioned as the former employer of AAI's CEO Sasha.

"Mobileye should be a very long-established company, always doing very strong edge processing or chip-level processing. They previously supplied chips to Tesla." — Tian Yuandong [00:51:22.920]

Meta (FAIR) Tian's former employer. Site of his research on Go (OpenGo) and his development of research taste through hands-on coding. Described as now suffering from organizational bloat that slows AI progress. Llama 4 and XDR cited as examples of teams that degraded past 150 people.

"Once Llama 4 and XDR exceeded 150 people they started falling apart. Beyond 150 people, it breaks down." — Tian Yuandong [01:22:05.800]

Google / Waymo Tian's entry point into the industry. He joined Google's self-driving car project in 2013. Noted for the repeated unfulfilled "5 years away" promises on autonomous vehicles.

"In 2011 I heard this thing — '5 years away.' Then I joined in 2013 — still '5 years away.' And now it might still be another 5 years." — Tian Yuandong [01:12:45.240]

Google DeepMind Mentioned for Alpha Evolve, which reduced the number of multiplications in matrix multiplication — an example of AI generating mathematically elegant results. Tian found it interesting but noted it was still "tree-level" (specific) rather than "Tao-level" (principled).

"Last year Google's Alpha Evolve achieved a reduction in the number of multiplications — reducing the count by some number. It rearranged the order of multiplications to save one addition or one multiplication, thereby reducing the cost. When you really go look at it, it's quite interesting." — Tian Yuandong [01:06:53.940]

Google Auto ML Google's ~10-year-old prior attempt at AI-optimizing AI. Tian explains why it failed: it lacked any model of high-level human knowledge, so humans had to manually define the search space, limiting generalization.

"The problem with that previous wave was mainly that we had no model that could represent human high-level knowledge and thinking. In that era, researchers defined the size and behavior of the search space, and AI was made into an algorithm to find better solutions within it." — Tian Yuandong [00:16:26.000]


4. People Identified

Tian Yuandong (田渊栋) Co-founder of Recursive Super Intelligence. Former principal researcher at Meta FAIR. Known for OpenGo (beat Korean professionals 20-0), work on efficiency of large models, Grokking, and theoretical understanding of neural networks. Has been building toward RSI since recognizing in August 2024 that GPT-5 could co-author a research paper with him.

"I used GPT-5 to co-author a paper — after finishing, I realized the model was already basically strong enough that I didn't need to hire a research assistant. Many ideas and approaches could be realized purely through dialogue with the model. This was when I thought: maybe in 5 years I'll be replaced. Rather than be replaced, I'll start a company." — Tian Yuandong [00:08:07.140]

Yann LeCun (乐坤) Chief AI Scientist at Meta. Cited as the archetypal "Type 1" researcher — someone who identifies the right direction (self-supervised learning, his famous cake diagram), but does not necessarily work through how to execute it. His conviction on self-supervised learning over reinforcement learning (because SSL has more feedback signal) turned out to be directionally correct and influenced GPT pretraining.

"LeCun, for instance — he wants to do self-supervised learning... He wouldn't go thinking about exactly how the supervision and feedback makes the model learn. He just grasps the point: more supervision is better than less, a model with more labels learns faster. He holds onto that and doesn't let go, persisting until people believe it has merit." — Tian Yuandong [01:16:42.380]

Ilya Sutskever Co-founder of OpenAI, now at SSI. Described as a "Type 1" researcher with near-religious conviction in Scaling Law. Tian credits him with being right at the right time, but notes even Ilya has acknowledged Scaling Law's room may be running out.

"Ilya, especially — his belief in Scaling Law is like a religious faith. He made people believe: don't overthink it, just scale, get the work done, put in the data and train. But I heard him say in an interview late last year that the room in Scaling Law is running out... I very much agree with that." — Tian Yuandong [01:18:39.780]

Dario Amodei CEO of Anthropic. Also categorized by Tian as a "Type 1" researcher — someone who points the direction rather than works through the implementation details.

"Ilya and Dario both belong to Type 1. They point the direction — especially Ilya, with his Scaling Law conviction." — Tian Yuandong [01:18:39.780]

Benham (Behnam Neyshabur) Researcher at Miraendar (the new RSI startup from DeepMind/Anthropic). Long-time theorist focused on understanding and analyzing models. Tian has known him for a long time and notes he frequently collaborated with Sanjeev Arora at Princeton.

"Benham — I've known him for a long time. He used to do theory, always working on model analysis and understanding, with many good papers. He often collaborated with Princeton's Sanjeev Arora. But many people, after this AI wave came, abandoned their prior work and just said OK, let's scale — and went to do other things." — Tian Yuandong [01:01:59.620]

Sanjeev Arora Princeton theoretical computer scientist, cited as a key collaborator in the AI theory community, mentioned in the context of researchers who have worked on understanding model behavior.

"He often collaborated with Princeton's Sanjeev Arora." — Tian Yuandong [01:02:29.260]

Jordan Keller Creator of the MUON optimizer. Initiated the NanoGPT Speedrun project over two years ago — a community benchmark that RSI's system has now achieved SOTA on.

"The GPT Speedrun is a project initiated more than two years ago by the developer of the MUON optimizer, Jordan Keller. Previously it was mainly humans doing it; now increasingly it's AI and large models doing it." — Host Manchi [01:27:28.600]

Sasha (AAI CEO, former Mobileye CEO) CEO of AAI, an Israeli semiconductor startup that competed against RSI on the SOL Exact Bench. Former Mobileye CEO. Represents the professional GPU/chip community that RSI (without GPU specialists) outperformed.

"Their CEO is Sasha, the former CEO of Mobileye. He's always done very strong edge processing or chip-level processing work. He's a professional semiconductor company. Their people are all professional GPU engineers — but our results were 10% better than theirs." — Tian Yuandong [00:51:22.920]

David (Amazon San Francisco Lab) Former head of Amazon's San Francisco AI Lab. Left to start a company; Tian doesn't know yet what they're building.

"Some Amazon people also left. The head of Amazon's San Francisco Lab, David, also came out. But what they're preparing to do now, I don't know." — Tian Yuandong [00:25:37.300]

Ma Yi (马亦) A researcher previously interviewed by the host. Noted for pursuing scientific understanding of AI and fundamental large model research, in contrast to pure scaling.

"Among people I've previously interviewed, I think Ma Yi is more in pursuit of this — the scientific understanding of things, and some fundamental large model research." — Host Manchi [01:03:28.780]


5. Operating Insights

The "Plateau-Then-Jump" Mental Model as a Strategic Planning Tool

For anyone building in AI — as an operator or investor — the most consequential strategic question is whether the capability curve is smooth or stepped. Tian's answer implies a specific portfolio and resource allocation logic: don't assume the current leader automatically compounds into dominance. Look for the next inflection point (i.e., the next capability that enables a new "S-curve"), and back teams positioned to discover it, not just teams scaling the current one.

"I think there are plateaus... It's a staged, leap-by-leap process. Are you believing in this being a straight-up line, or does it have plateaus? I think it has plateaus." — Tian Yuandong [00:40:12.400]

150-Person Team Size Is a Hard Ceiling for Frontier AI Research Effectiveness

Operators scaling AI research organizations should treat ~150 people as an inflection point where the two-tier (thinker/executor) dynamic begins to dominate and research velocity drops. Small, flat, mutually familiar teams where everyone is hands-on are structurally faster. This is actionable for both lab leaders and investors evaluating team structures.

"Once Llama 4 and XDR exceeded 150 people they started falling apart. With a private team of 30, everyone knows everyone — relationships are very simple. Once more people join, and you have middle management, the problem starts again." — Tian Yuandong [01:22:05.800]

Use RSI's Staircase Structure to Build Commercially Viable Milestones Into Your Roadmap

Unlike an all-or-nothing moonshot, RSI produces commercially valuable output at each stage. Teams building toward RSI should design their product roadmap around each "step" delivering standalone value (e.g., kernel optimization, training efficiency) rather than waiting for the ultimate vision. This enables revenue and credibility-building in parallel with long-horizon research.

"Even if we haven't achieved the hardest goal, this thing already has many applications. The first step itself already has a lot of practical value. The second step will have even deeper things to excavate." — Tian Yuandong [00:36:45.520]


6. Overlooked Insights

GLM 5.2 and Open-Source Models Are Now Functionally Close to Frontier Closed Models for Execution Tasks — With Specific Failure Modes Documented

Tian briefly tests GLM 5.2 on return to China and gives a nuanced on-the-record technical assessment: it handles execution tasks well but can get stuck in infinite loops on harder reasoning tasks. This throwaway comment is actually significant. If a leading frontier researcher with direct experience using Anthropic's Claude 4.6/4.7/4.8 (and who finds those versions barely distinguishable from each other) says the primary difference with Claude is now context window length rather than capability, and that a Chinese open-source model is "pretty good" for execution — this suggests the capability gap between frontier closed models and leading open-source models may be closing faster than market valuations reflect.

"GLM 5.2 is still better at doing execution-type work, and it does that well. But sometimes it may fall into infinite loops, or have various rough edges. Overall though, it's quite good... Compared to Opus 4.7, I think GLM 5.2 is quite good. As for 4.6, 4.7, 4.8 — I don't feel a huge difference between them. Maybe the only real difference with 4.8 is that the context window is longer." — Tian Yuandong [00:29:03.140]

The RSI Data Problem Is Totally Unsolved — and the Difficulty/Quality Tradeoff Is the Actual Bottleneck

In a brief exchange that passes quickly, Tian reveals that the data strategy for training RSI systems is entirely open and team-internal, with no consensus in the field. More specifically, he signals that the easy data methods produce poor results, and the effective methods are expensive or hard — making this the actual rate-limiting factor, ahead of compute or model architecture. This means whoever cracks the RSI data flywheel first has a durable moat that is not visible from published benchmarks or model releases.

"There's no consensus yet. There are many methods for making this data — the specific methods, logic, principles, and approaches are not yet clear. Various methods are being tried. Everyone is still in exploration mode... Some methods are easy, some are hard. The problem between hard and easy is: what effect do they actually have? It's a tradeoff. The hard ones probably have good effects. The easy ones probably don't." — Tian Yuandong [00:55:09.800]