Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/THEMES/REINFORCEMENT LEARNING
// THEME

Reinforcement Learning

COMPANIES 39VELOCITY — STABLECAPITAL 28D $37064.9M · 49 DEALS
TOP INVESTORS: nvidia (34) · google (9) · khosla ventures (6) · openai (6) · sequoia (6)

CAPITAL FIGURES ARE MEDIA-EXTRACTED ESTIMATES, NOT VERIFIED FILINGS.

Mega-rounds dominate; weekly capital remains volatile
$18.4B · wk of 07-13 ▶2026-05-04 ── 2026-07-27 · WEEKLY
Strategic and Series C rounds command lion's share of capital
unknown
$38.0B · 67 DEALS
series a
$11.6B · 19 DEALS
series c
$8.8B · 11 DEALS
seed
$3.0B · 10 DEALS
series d plus
$10.5B · 8 DEALS
Mention momentum
MENTIONS / WEEK · PEAK 302

EXTRACTED FROM 25+ PODCASTS & VC NEWSLETTERS · MEDIA-REPORTED FIGURES, NOT VERIFIED FILINGS

// THE LEAD
▲ STRENGTHENING

VLA foundation models become the RL backbone for physical AI

Vision-language-action models are cementing their role as the dominant RL substrate for physical AI, with Physical Intelligence commanding a $27.5B valuation and $2.5B Series C backed by Nvidia, Sequoia, Lightspeed, JPMorgan, and B Capital. Benchmarks are sharpening the competitive picture: Qwen-RobotManip has claimed the top rank on RoboChallenge with a 20% relative improvement over π0.5 in out-of-distribution settings, while RLWRLD's RLDX-1 targets dexterity-first industrial manipulation across humanoid and single-arm embodiments. The VLA paradigm is no longer a research curiosity — it is the commercial architecture through which capital is concentrating at scale.

// TRENDS
▲ STRENGTHENINGAutonomous RL self-improvement loops arrive in physical robotics

The frontier of physical AI RL has shifted from human-supervised training to autonomous policy self-improvement. The ENPIRE research framework enables robots to iteratively refine their own policies without human intervention, while Prime Intellect has run large-scale autonomous AI research experiments that demonstrate RL loops operating without per-step human oversight. This structural shift compresses the cost and time to deploy capable robot policies across novel environments.

Why it matters · Autonomous self-improvement loops break the data-labeling bottleneck that has constrained robotics RL, unlocking exponential capability scaling for operators who adopt these frameworks early.

▲ STRENGTHENINGSimulation-to-real transfer matures into investable infrastructure

Sim-to-real pipelines have crossed from research tooling into production-grade infrastructure. NVIDIA's Isaac Gym platform is running 62,000 parallel environments on RTX 5090 GPUs using SAPG optimization, compressing locomotion policy retraining to two hours on a single GPU — a task previously scoped at weeks. Applied Compute and General Intuition are building complementary infrastructure layers that convert game-derived and synthetic environments into real-world-transferable robot and agent policies.

Why it matters · As sim-to-real becomes reliable and cheap, the marginal cost of training diverse robot skills collapses — accelerating time-to-deployment for hardware companies and creating durable moats for simulation platform providers.

▲ STRENGTHENINGNVIDIA consolidates RL infrastructure dominance across the stack

With 28 deals in the top-investor rankings — more than triple Amazon's count — NVIDIA is the most active strategic backer across the RL ecosystem, appearing in the Physical Intelligence Series C ($2.5B), the $800M Series C alongside General Catalyst and Vista Equity, and a $200M seed round alongside Kleiner Perkins and a16z. Isaac Gym's GPU simulation stack, cited as the compute backbone for frontier RL experiments, reinforces hardware lock-in at the infrastructure layer.

Why it matters · NVIDIA's dual role as chip supplier and lead investor creates compounding competitive advantages: portfolio companies build on NVIDIA silicon and software, deepening platform dependency across the RL value chain.

▲ NEWAgentic RL environments emerge as a distinct enterprise software category

Deeptune is pioneering a new product archetype — high-fidelity RL training environments that simulate enterprise software workflows across tools like Slack and Salesforce — signaling that the next frontier for reinforcement learning is white-collar task automation rather than physical robotics alone. Signal [49] captures the market-level shift: the highest-leverage AI skill has moved from prompt engineering to architecting repeating agent loops, exactly the capability Deeptune's environments are designed to train.

Why it matters · Enterprise agentic RL environments address a multi-trillion-dollar automation market and represent a defensible infrastructure layer between frontier model providers and enterprise deployments.

StrictlyVC · Jul 9The AI Corner · Jul 7Data Driven VC · Jul 9
// COMPANIES
39 COMPANIES
01
Anthropic
anthropic.com
GROWTH · JUL 29
982 SIGNALS · LAST SEEN AUG 2, 2026
02
OpenAI
openai.com
812 SIGNALS · LAST SEEN AUG 2, 2026
03
Nvidia
nvidia.com
510 SIGNALS · LAST SEEN AUG 2, 2026
04
Google
google.com
$830M · SERIES A · SITUATIONAL AWARENESS + BLACKROCK · JUL 27
321 SIGNALS · LAST SEEN JUL 31, 2026
05
Fireworks
fireworks.ai
$1.5B · GROWTH · INDEX VENTURES + GAVIN BAKER · JUL 23
45 SIGNALS · LAST SEEN AUG 2, 2026
06
Applied Compute
appliedcompute.com
$20M · UNKNOWN · KBR + DATABRICKS VENTURES · JUL 23
13 SIGNALS · LAST SEEN JUL 23, 2026
07
Physical Intelligence
physicalintelligence.company
$5.0B · SERIES A · JUL 16
126 SIGNALS · LAST SEEN JUL 29, 2026
08
Prime Intellect
primeintellect.ai
7 SIGNALS · LAST SEEN JUL 14, 2026
09
Tencent
tencent.com
$150M · PRE-SERIES B · JUL 7
50 SIGNALS · LAST SEEN JUL 31, 2026
10
General Intuition
$320M · SERIES A · JUL 2
10 SIGNALS · LAST SEEN JUL 19, 2026
11
Mirondale
1 SIGNAL · LAST SEEN JUL 1, 2026
12
EquiLibre Technologies
SERIES A · CREANDUM · JUL 1
2 SIGNALS · LAST SEEN JUL 1, 2026
13
Agility Robotics
agilityrobotics.com
SPAC · JUN 25
26 SIGNALS · LAST SEEN JUL 28, 2026
14
Google DeepMind
deepmind.google
$75M · GOOGLE DEEPMIND · JUN 23
62 SIGNALS · LAST SEEN JUL 31, 2026
15
Mercor
mercor.com
GROWTH · JUN 3
27 SIGNALS · LAST SEEN JUL 25, 2026
16
Applied Intuition
appliedintuition.com
SERIES A · LUX CAPITAL + A16Z (MARC ANDREESEN) · MAY 31
11 SIGNALS · LAST SEEN JUL 21, 2026
17
Parallel
parallel.com
$0M · SEED · KHOSLA VENTURES + CHARLOTTE INDEX · MAY 31
4 SIGNALS · LAST SEEN MAY 31, 2026
18
Deeptune
deeptune.ai
$0M · SERIES A · A16Z + FELICIS VENTURES · MAY 31
3 SIGNALS · LAST SEEN JUN 1, 2026
19
Recursive
4 SIGNALS · LAST SEEN JUL 30, 2026
20
Google DeepMind
deepmind.com
94 SIGNALS · LAST SEEN JUL 29, 2026
21
Core Automation
6 SIGNALS · LAST SEEN JUL 29, 2026
22
arXiv
arxiv.org
17 SIGNALS · LAST SEEN JUL 28, 2026
23
Carnegie Mellon University
cmu.edu
33 SIGNALS · LAST SEEN JUL 28, 2026
24
arXiv Physical AI
10 SIGNALS · LAST SEEN JUL 27, 2026
25
DeepMind
deepmind.com
24 SIGNALS · LAST SEEN JUL 27, 2026
26
Midea Group
midea.com
12 SIGNALS · LAST SEEN JUL 26, 2026
27
NeoteAI
neoteai.com
17 SIGNALS · LAST SEEN JUL 26, 2026
28
Freesolo Flash
1 SIGNAL · LAST SEEN JUL 24, 2026
29
HIVE Robots
4 SIGNALS · LAST SEEN JUL 22, 2026
30
Turing
turing.com
3 SIGNALS · LAST SEEN JUL 22, 2026
31
TU Darmstadt
3 SIGNALS · LAST SEEN JUL 21, 2026
32
Zioneer Robot Team
2 SIGNALS · LAST SEEN JUL 3, 2026
33
UC Berkeley
berkeley.edu
37 SIGNALS · LAST SEEN JUN 30, 2026
34
University of North Carolina at Chapel Hill
6 SIGNALS · LAST SEEN JUN 26, 2026
35
ENPIRE
2 SIGNALS · LAST SEEN JUN 18, 2026
36
CMU
1 SIGNAL · LAST SEEN JUN 18, 2026
37
Ideogram
ideogram.ai
10 SIGNALS · LAST SEEN JUN 15, 2026
38
SenseNova
sensenova.sensetime.com
14 SIGNALS · LAST SEEN JUN 15, 2026
39
RLWRLD
rlwrld.ai
16 SIGNALS · LAST SEEN MAY 15, 2026