Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/THEMES/REINFORCEMENT LEARNING
// THEME

Reinforcement Learning

COMPANIES 49VELOCITY ▼ COOLINGCAPITAL 28D $0.0M · 1 DEAL

CAPITAL FIGURES ARE MEDIA-EXTRACTED ESTIMATES, NOT VERIFIED FILINGS.

Mention momentum
MENTIONS / WEEK · PEAK 24

EXTRACTED FROM 25+ PODCASTS & VC NEWSLETTERS · MEDIA-REPORTED FIGURES, NOT VERIFIED FILINGS

// THE LEAD
▲ STRENGTHENING

Physical AI RL closes the gap on frontier labs

Temporal GRPO, published on arXiv Physical AI, now outperforms Physical Intelligence's π0 by 26.6 percentage points (75.8% vs 49.2%) on the RoboTwin 2.0 benchmark, signaling that open research is rapidly closing the gap with well-funded proprietary VLA systems. The core technical breakthrough — solving trajectory-level credit aliasing, where valid early actions are penalized by later failures — unlocks far more sample-efficient robot policy training. Google DeepMind's concurrent release of Gemini Robotics 2 and NVIDIA's GR00T N1 generalist baseline confirm that physical RL is now a multi-front arms race among the world's best-resourced labs, with Temporal GRPO methods explicitly applicable to post-training GR00T models.

// TRENDS
▲ STRENGTHENINGNVIDIA consolidates RL infrastructure dominance across the stack

NVIDIA appears in the two largest disclosed rounds of the period — a $2B growth round (alongside Blackstone, Jane Street, and Coatue, at a $10.5B valuation) and a $1.1B round alongside AMD Ventures and General Catalyst — while its hardware (RTX PRO 6000, Jetson Thor, RTX 3090/4090/5090) underpins all major physical AI benchmarks. The open-sourcing of Cosmos 3, an omni-modal world foundation model combining video, audio, language, and action signals, extends NVIDIA's stack from chips into training frameworks, synthetic data, and model weights — making it the de facto RL infrastructure vendor from silicon to simulation.

Why it matters · NVIDIA's platform lock-in strategy means that rivals building RL stacks on commodity compute will face a compounding disadvantage as Cosmos-trained models outperform alternatives at every layer.

▲ STRENGTHENINGAgentic RL environments emerge as a distinct enterprise software category

Deeptune's high-fidelity RL environments that simulate multi-step workplace workflows across Slack and Salesforce, and Freesolo Flash's commodity RL fine-tuning for small language models, represent a new category of enterprise software that sits below the model layer. The Anthropic IPO announcement — targeting October 2026 — and Anthropic's $6B acquisition of Decart both inject fresh urgency into the agent training infrastructure market, as enterprises will need purpose-built environments to fine-tune and evaluate agentic systems before deployment.

Why it matters · Enterprise software buyers who control the simulation and evaluation layer will capture a recurring, defensible revenue stream as every major company begins deploying RL-trained agents into production workflows.

CORROBORATED · 2 SOURCE TYPES20VC · Aug 13Axios Pro Rata · Aug 13The AI Corner · Aug 12
▲ STRENGTHENINGAutonomous RL self-improvement loops arrive in physical robotics

The ENPIRE research framework enables autonomous robot policy self-improvement without human intervention, while RLWRLD's RLDX-1 foundation model integrates vision, force sensing, and memory across single-arm, dual-arm, and humanoid embodiments. Prime Intellect's large-scale autonomous AI research experiments and General Intuition's spatiotemporal reasoning models trained on game-play video extend the self-improvement paradigm beyond physical robots into 3D agentic environments — together suggesting that closed-loop RL training is becoming a systems engineering problem rather than a research one.

Why it matters · Once self-improvement loops are productized, the cost of capability gains drops sharply, accelerating the timeline to commercially deployable humanoid robots and compressing the investment window for early bets.

▲ NEWTop RL talent exodus from big tech seeds next generation of labs

Jeff Dean's departure from Google to found a science-focused AI lab — reportedly co-led with Vinod Khosla of Khosla Ventures — mirrors the pattern that produced OpenAI and Anthropic, with Khosla described as 'doing the playbook again.' Multiple senior departures from OpenAI ahead of its IPO, including safety-focused researchers, are simultaneously seeding new ventures. Applied Intuition's autonomous vehicle simulation business and the founding signals around new compute-focused labs indicate that the RL talent pool is now diffusing from three or four anchor institutions into a broader ecosystem of specialized labs.

Why it matters · Investors who backed Khosla at the OpenAI and Anthropic stages should treat the Jeff Dean lab formation as a structural signal — the next generation of RL frontier labs is being seeded now, before products exist.

// COMPANIES
49 COMPANIES
01
Google DeepMind
deepmind.com
SERIES A · AUG 22
132 SIGNALS · LAST SEEN SEP 9, 2026
02
Deeptune
deeptune.ai
$0M · SERIES A · A16Z + FELICIS VENTURES · MAY 31
3 SIGNALS · LAST SEEN JUN 1, 2026
03
Toyota Research Institute
tri.global
GIFT/GRANT · TOYOTA RESEARCH INSTITUTE + FLEXIV · APR 30
7 SIGNALS · LAST SEEN AUG 26, 2026
04
Harbin Institute of Technology
hit.edu.cn
9 SIGNALS · LAST SEEN SEP 18, 2026
05
DeepMind
deepmind.com
35 SIGNALS · LAST SEEN SEP 18, 2026
06
UC Berkeley
berkeley.edu
38 SIGNALS · LAST SEEN SEP 14, 2026
07
Carnegie Mellon University
cmu.edu
43 SIGNALS · LAST SEEN SEP 13, 2026
08
ETH Zurich
ethz.ch
4 SIGNALS · LAST SEEN SEP 9, 2026
09
LeRobot
2 SIGNALS · LAST SEEN SEP 5, 2026
10
DeepMind
deepmind.google
2 SIGNALS · LAST SEEN SEP 3, 2026
11
Peking University
pku.edu.cn
19 SIGNALS · LAST SEEN SEP 3, 2026
12
Shanghai Qi Zhi Institute
sqz.ac.cn
1 SIGNAL · LAST SEEN AUG 26, 2026
13
Tsinghua University
tsinghua.edu.cn
38 SIGNALS · LAST SEEN AUG 26, 2026
14
CMU
1 SIGNAL · LAST SEEN AUG 24, 2026
15
ICML
1 SIGNAL · LAST SEEN AUG 19, 2026
16
Leela Chess Zero
1 SIGNAL · LAST SEEN AUG 13, 2026
17
Farama Foundation
farama.org
2 SIGNALS · LAST SEEN AUG 12, 2026
18
RLBench
3 SIGNALS · LAST SEEN AUG 4, 2026
19
ManiSkill3
2 SIGNALS · LAST SEEN AUG 4, 2026
20
TU Darmstadt
3 SIGNALS · LAST SEEN JUL 21, 2026
21
UT Austin
utexas.edu
1 SIGNAL · LAST SEEN JUL 16, 2026
22
MIT CSAIL
csail.mit.edu
8 SIGNALS · LAST SEEN JUN 25, 2026
23
Cadence Design Systems
cadence.com
4 SIGNALS · LAST SEEN JUN 18, 2026
24
Mila
mila.quebec
2 SIGNALS · LAST SEEN JUN 11, 2026
25
AIRe Lab
1 SIGNAL · LAST SEEN JUN 10, 2026
26
Simon Fraser University
sfu.ca
1 SIGNAL · LAST SEEN JUN 2, 2026
27
University of Michigan
umich.edu
2 SIGNALS · LAST SEEN JUN 1, 2026
28
RoboTwin
1 SIGNAL · LAST SEEN JUN 1, 2026
29
TU Delft
0 SIGNALS · LAST SEEN MAY 29, 2026
30
SERL
1 SIGNAL · LAST SEEN MAY 28, 2026
31
Mujoco
0 SIGNALS · LAST SEEN MAY 21, 2026
32
Great Bay University
gbu.edu.cn
4 SIGNALS · LAST SEEN MAY 18, 2026
33
Nankai University
nankai.edu.cn
2 SIGNALS · LAST SEEN MAY 18, 2026
34
University College London
ucl.ac.uk
6 SIGNALS · LAST SEEN MAY 15, 2026
35
University of Oxford
ox.ac.uk
2 SIGNALS · LAST SEEN MAY 15, 2026
36
University of Pennsylvania
upenn.edu
2 SIGNALS · LAST SEEN MAY 5, 2026
37
HKUST (Guangzhou)
hkust-gz.edu.cn
1 SIGNAL · LAST SEEN MAY 4, 2026
38
Stanford Artificial Intelligence Laboratory (SAIL)
0 SIGNALS · LAST SEEN JAN 1, 2024
39
Preferred Networks, Inc.
0 SIGNALS · LAST SEEN JAN 1, 2024
40
Robotics and Embodied AI Lab @ UdeM/MILA
0 SIGNALS · LAST SEEN JAN 1, 2024
41
RLLab McGill/Mila
0 SIGNALS · LAST SEEN JAN 1, 2024
42
Nirvanic | Quantum minds to make robots work
0 SIGNALS · LAST SEEN JAN 1, 2024
43
Learning Systems and Robotics Lab
0 SIGNALS · LAST SEEN JAN 1, 2024
44
Alberta Machine Intelligence Institute (Amii)
0 SIGNALS · LAST SEEN JAN 1, 2024
45
Mila - Institut québécois d'intelligence artificielle
0 SIGNALS · LAST SEEN JAN 1, 2024
46
International Max Planck Research School for Intelligent Systems (IMPRS-IS)
0 SIGNALS · LAST SEEN JAN 1, 2024
47
Karlsruhe Institute of Technology (KIT)
0 SIGNALS · LAST SEEN JAN 1, 2024
48
Massachusetts Institute of Technology
0 SIGNALS · LAST SEEN JAN 1, 2024
49
Mila - Quebec Artificial Intelligence Institute
0 SIGNALS · LAST SEEN JAN 1, 2024