Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/JACK CLARK FROM IMPORT AI/Import AI 469: Science AI; RSI s…
NEWS
// NEWSLETTER ISSUE
JACK CLARK FROM IMPORT AI

Import AI 469: Science AI; RSI simulator; and Zuck's technological pessimism

DATE August 17, 2026SOURCE JACK CLARK FROM IMPORT AIPARTICIPANTS JACK CLARK FROM IMPORT AI
// SUMMARY

1. Key Themes


Theme 1: AI "Scientific Taste" as the New Capability Frontier

The emerging benchmark for AI is no longer raw performance but the ability to autonomously discover rules, form hypotheses, and conduct research — what the article calls "taste."

"The new frontier for analyzing AI systems is understanding how good they are at inferring the unwritten rules of their environment."

"Tests like this are attempts to isolate a prerequisite for creativity, which is being able to autonomously discover useful undocumented things about novel situations you find yourself in."


Theme 2: Recursive Self-Improvement (RSI) Is the Horizon Event to Watch

Multiple threads in the newsletter converge on RSI as the pivotal near-term inflection point, with current research specifically building toward that threshold.

"My guess is we'll reach human parity on DiG-bench by middle of 2027, at which point we should expect things like recursive self-improvement to seriously kick off."

"The better AI systems get at this, the more likelihood we can assign to the idea that AI systems will imminently become capable of building themselves."


Theme 3: Small Supervisory Models Layered Over Frontier Models as an Architectural Pattern

Inherent's Faraday demonstrates a powerful and investable pattern: a small, purpose-trained model acting as a "scientific director" over larger, general-purpose frontier models.

"The company built a supervisory harness and relatively small LLM which sits on top of large, proprietary frontier models, and controls them in a way that improves their effectiveness at science."

"The skills Faraday acquires – deciding what to investigate, scoping experiments to a budget, and judging a replication – compound with advances in frontier coding models. One might hope that a single post-trained outer agent can track the frontier as better models are released, at least over some time period."


Theme 4: Democratization vs. Concentration of Superintelligence as the Defining Policy Debate

Meta's strategic bet is mass proliferation of AI tools as a safety mechanism. This framing will shape regulatory, competitive, and investment landscapes.

"The defining questions of our age are who will have access to superintelligence and what will we direct it towards. We propose a philosophy based on individual empowerment as the source of prosperity, invention as the primary purpose of superintelligence, and balance of power as the foundation of safety."


2. Contrarian Perspectives


Perspective 1: Zuckerberg's proliferation strategy contains a fatal logical gap The consensus framing of Meta's open AI strategy is that democratizing access prevents dangerous power concentration. Clark argues this ignores the agency problem: a system capable of superhuman invention may not remain a neutral tool in service of individual empowerment.

"The missing question in all of this is 'will a system capable of superhuman invention solely work on behalf of the individual empowerment of people that are less capable than it at invention?' Surely this is the key question?"

"I am not suggesting that superhuman invention guarantees some kind of malign entity that is independent from people. Rather I am suggesting that it's hard to reconcile a system capable of superhuman invention with something that doesn't fundamentally alter the balance of power in the world in ways that are confusing and hard to reason about."


Perspective 2: The 20% human-tier benchmark score is more alarming than reassuring The surface read of DiG-bench results is that AI still lags humans badly. Clark inverts this: a 20% success rate on Tier 7 tasks — which individual humans solve at 100% — already signals meaningful, near-term discovery capability.

"Only Opus 5 and Fable 5 were able to beat any tasks (0.2) in Tier 7... a 20% success rate on Tier 7 is pretty poor compared to the fact individual humans were able to get 100% on the tests."

"My guess is we'll reach human parity on DiG-bench by middle of 2027, at which point we should expect things like recursive self-improvement to seriously kick off."


Perspective 3: A small 27B post-trained model can outperform frontier giants on specialized scientific tasks Conventional wisdom assumes bigger frontier models dominate. Faraday, a 27B model, beats flagship-class models (Opus 4.8, GPT-5.5) on the majority of research replication tasks — suggesting domain-specialized small models may be structurally superior for agentic science workflows.

"Faraday using Codex is able to beat standard Opus 4.8 and GPT-5.5 on some replication tasks, exceeding their performance 'on 73% of in-distribution ML tasks, and on 60% of held-out AI-for-science tasks, according to our rubric-based judge.'"


3. Companies Identified


Inherent

  • Description: AI startup building Faraday, an autonomous AI scientist
  • Why mentioned: Case study in building supervisory LLM agents that outperform frontier models on scientific research tasks
  • Quote: "Researchers with AI startup Inherent have published a paper showing how they are building Faraday, an AI scientist model that they hope can develop some taste in terms of research."

Paradigm Research

  • Description: Research organization building educational simulations about AI development dynamics
  • Why mentioned: Creator of the RSI Simulator, a browser-based game modeling recursive self-improvement economics
  • Quote: "Here's a fun game from the folks at Paradigm Research which aims to simulate what it's like to run a company building AI systems which become capable of recursive self-improvement."

Meta

  • Description: Global technology conglomerate, developer of open-weight AI models
  • Why mentioned: Zuckerberg's essay lays out Meta's strategic philosophy for superintelligence — democratized access as a safety strategy — which Clark subjects to critical scrutiny
  • Quote: "Mark Zuckerberg has written an essay called 'The Future is for Everyone' that serves as something of a manifesto for how he and Meta are approaching the development of AI systems."

OpenAI

  • Description: Leading AI lab
  • Why mentioned: OpenAI's Codex is used as the underlying coding agent tool inside Faraday; GPT-5.5 is a benchmark comparator in DiG-bench
  • Quote: "Faraday is a 27B model that uses a coding agent (OpenAI Codex) as an underlying tool."

Anthropic

  • Description: AI safety-focused frontier lab
  • Why mentioned: Claude Opus 5 is a top performer on DiG-bench; Claude Opus 4.7 is used as a grader in the Replica dataset construction; Opus 4.8 is a benchmark comparator for Faraday
  • Quote: "Opus 5 and Fable 5 with Claude Code are the best overall models, followed by GPT-5.5."

4. People Identified


Jürgen Schmidhuber

  • Description: Pioneer AI researcher; co-inventor of LSTM; founder of AI lab NNAISENSE
  • Why mentioned: Co-author of the DiG-bench paper, cited as creative credibility signal for the research
  • Quote: "One of the authors is Juergen Schmidhuber, an extremely creative OG AI researcher."

Mark Zuckerberg

  • Description: CEO of Meta
  • Why mentioned: Author of "The Future is for Everyone," a manifesto on Meta's AI philosophy; Clark uses it as a foil to argue that superintelligence proliferation doesn't automatically produce safe or equitable outcomes
  • Quote: "The defining questions of our age are who will have access to superintelligence and what will we direct it towards."

Jack Clark

  • Description: Co-founder of Anthropic; author of Import AI newsletter
  • Why mentioned: Author and analyst offering the critical framing throughout the piece, including the RSI timeline prediction and the Zuckerberg critique
  • Quote: "My guess is we'll reach human parity on DiG-bench by middle of 2027, at which point we should expect things like recursive self-improvement to seriously kick off."

5. Operating Insights


Insight 1: Build evaluation datasets by removing results from existing papers, not by generating novel benchmarks from scratch Inherent's approach to constructing Replica is a replicable, scalable methodology for any team building AI research agents: use real published work as ground truth, systematically remove key results, and train/evaluate on the agent's ability to reconstruct them. This is cheaper and more rigorous than synthetic data generation.

"They assemble a dataset ('Replica') consisting of research papers that have key graphs or results missing from them, then they see how well AI systems can autonomously do experiments that fill in the blanks."


Insight 2: When building agentic AI products, design your outer agent to be model-agnostic so it "tracks the frontier" as underlying models improve Faraday's architecture deliberately separates the supervisory reasoning layer from the raw coding/execution layer, allowing it to benefit automatically from improvements in foundation models without retraining.

"One might hope that a single post-trained outer agent can track the frontier as better models are released, at least over some time period."


Insight 3: Use game-based simulations to build organizational intuition about complex AI dynamics before making resource allocation decisions The RSI Simulator is not just a consumer toy — it models resource tradeoffs (researchers vs. compute, data licensing timing) that are directly relevant to AI lab and product strategy decisions.

"If you play the game you can get a good feel for how different components of AI research interact, ranging from how you balance investing in researchers versus compute, how and when to license data, and more."


6. Overlooked Insights


Insight 1: Private benchmark holdback as a structural anti-contamination method is becoming standard practice DiG-bench's decision to keep the majority of games private — released only after evaluation — is a methodological signal about how serious benchmark designers are addressing training data contamination. Investors and builders evaluating AI capability claims should weight benchmarks that use holdout-private evaluation far more heavily than public ones.

"All of these games have been built by human experts. The majority of the games are kept private so that AI systems don't train on them."


Insight 2: Per-turn credit assignment in multi-step agent training is an underappreciated technical lever Inherent's Faraday uses a judge model to provide not just an overall reward but "per-turn credit assignment weights" — a granular training signal that helps the agent learn which specific steps in a long research workflow contributed to success or failure. This is a meaningful advancement over binary success/failure RL that most teams default to.

"A Codex-based Judge model... provides an overall reward and per-turn credit assignment weights, which are used to train the Faraday agent using a modified version of GRPO."