Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/GUIDES/VLA MODELS
GUIDE

VLA Models Explained: Vision-Language-Action, the Architecture Behind Robot Brains (2026)

VLA models take in camera images and a sentence, and output motor actions. How they work, the lineage from RT-2 to π and Helix, and where to follow the research as it lands.

Bryan Altman
Bryan Altman
Founder, Teahose · angel investor & builder
Updated 2026-06-23

Key takeaways

  • A VLA (vision-language-action) model is one neural network that takes in camera images plus a plain-language instruction and outputs robot motor actions — it is the architecture behind nearly every "robot foundation model" of the last two years.
  • VLA work is one of the most-discussed research threads in our corpus: vision-language-action models surface in 60 of the 1,150+ expert podcast, newsletter, and research summaries we have analyzed.
  • The lineage to know runs RT-2 → OpenVLA → Physical Intelligence's π series → Figure's Helix → Gemini Robotics and Nvidia's GR00T, all converging on bigger pretrained backbone + more diverse robot data + a faster action decoder.
  • Most coverage chases funding headlines; what actually predicts the sector is the research, because rounds lag the papers by months — so track the papers, not the press releases.

Share of voice: the companies this guide covers, by mentions across Teahose's 1,150+ expert AI conversations
Share of voice: the companies this guide covers, by mentions across Teahose's 1,150+ expert AI conversations

Each bar counts how many of Teahose's 1,150+ expert summaries mention it (word-boundary match across our podcast, newsletter, and paper corpus, June 2026).

Track the field: find the companies most similar to Physical Intelligence and get their latest funding and product signals by email — Teahose Lookalikes.

Mention counts from Teahose's analysis of 1,150+ expert podcast, newsletter & research summaries, June 2026.

A vision-language-action (VLA) model is a neural network that turns what a robot sees, plus an instruction in plain language, into motor actions. It's the architecture behind essentially every "robot foundation model" headline of the last two years — and the reason robotics suddenly sounds like the early LLM era.

The Idea in One Paragraph

Take a vision-language model — the kind that can look at a photo and answer questions about it — and teach it one more output type: actions. The VLM already knows what a "mug" is, what "left of the sink" means, and roughly how objects behave, from internet pretraining. Fine-tune it on robot demonstration data (camera frames + instructions + the motions that succeeded) and that knowledge transfers to control. The robot generalizes not because it practiced every task, but because its backbone has seen the world.

That transfer is the entire economic story: it converts robotics from a per-task engineering business into a data-and-scale business — the shape of business venture capital knows how to price. (The pricing itself: see our Physical Intelligence, Skild, and Figure valuation guides.)

The Lineage That Matters

  • RT-2 (Google DeepMind, 2023) — the proof of concept: actions tokenized like text, emitted by a VLM. Manipulation success on novel objects roughly doubled vs the non-pretrained baseline.
  • OpenVLA (2024) — the open-source moment: a 7B VLA trained on the multi-robot Open X-Embodiment data, competitive with closed models; the reference point academic work builds on.
  • π series (Physical Intelligence, 2024→) — the commercial flagship: π0 introduced flow-matching action generation for smooth, high-frequency control across many robot bodies; successive π releases are the benchmark others quote.
  • Helix (Figure, 2025→) — the humanoid deployment: a fast-slow split (a large VLM reasons slowly; a small policy acts at high frequency) running Figure's robots at BMW — the architecture most teams converged toward.
  • Gemini Robotics (Google DeepMind, 2025→) and GR00T (Nvidia) — frontier-lab and ecosystem plays at the same layer.

The pattern across all of them: bigger pretrained backbone + more diverse robot data + a faster action decoder = more general behavior. Whether that curve keeps bending — or hits a data wall, since robot demonstrations are vastly scarcer than text — is the live scientific question of physical AI.

What VLAs Still Don't Solve

Reliability at deployment grade (a 95%-success policy fails one pick in twenty — unacceptable on a production line), long-horizon tasks (VLAs act; they barely plan), and data hunger (every lab is betting on a different mix of teleoperation, simulation, and video learning to feed the models). These gaps are exactly what the current paper wave attacks, which is why following the research is the leading indicator for the whole sector — funding rounds lag the research results by months.

Is VLA Already Dead? The World-Model Debate

A real 2026 argument worth knowing before you bet the term: some researchers — most loudly Nvidia's Jim Fan — argue the pure VLA paradigm is a stepping stone, and that "world-action models" (predict the next world state, then choose actions inside that learned model, refined with reinforcement learning) will supersede it, mirroring how language models evolved beyond next-token-only training. The counter-case is empirical: VLAs are what actually ships today — Helix runs Figure's humanoids on a real BMW line, the π series is deployed across robot bodies — while world models remain mostly research. The likely resolution is convergence, not replacement: the hottest current thread pairs a VLA policy with a world model that trains and evaluates it in imagination. So "VLA is dead" is best read as "the action head won't stay the whole story," not as a reason to stop tracking the architecture. Either way, the signal is the same: watch the papers, where this gets decided, not the press releases.

The Companies Building This Layer

Live from the Teahose intel graph

VLA & Robot Foundation Model Companies by Signal Volume

Live membership of the vision-language-action-models and embodied-foundation-models themes · ranked by extracted signals

  1. 01Nvidialast seen JUL 23461 signals
  2. 02Googlelast seen JUL 24289 signals
  3. 03Physical Intelligencelast seen JUL 23122 signals
  4. 04Google DeepMindlast seen JUL 2193 signals
  5. 05Stanford Universitylast seen JUL 2178 signals
  6. 06Google DeepMindlast seen JUL 2255 signals
  7. 07Hugging Facelast seen JUL 2454 signals
  8. 08Figurelast seen JUL 2446 signals
  9. 09Tencentlast seen JUL 2045 signals
  10. 10UC Berkeleylast seen JUN 3037 signals
  11. 11Samsunglast seen JUL 2429 signals
  12. 12Project Prometheuslast seen JUL 229 signals
  13. 13Physical Intelligencelast seen JUL 1727 signals
  14. 14Xiaomi Roboticslast seen JUL 2324 signals
  15. 15Generalist AIlast seen JUL 1418 signals
Updated continuously as new signals landExplore the full VLA theme

Follow the Research, Not the Press Releases

Our papers pipeline scores new robotics and AI research daily and publishes summaries of the top one or two — VLA, manipulation, and world-model work dominates. It's the cheapest way to see this field move before the funding announcements do.

Related: What is physical AI? · Physical Intelligence valuation · Figure AI valuation · robotics startups, live-ranked · the world models theme.

Bottom line: A VLA model is the single network that turns a robot's camera images plus a plain-language instruction into motor actions, and it's the architecture behind nearly every robot foundation model of the last two years — to see where it goes next, track the research, because funding rounds lag the papers by months.

Concepts are stable; the model lineage and live map are as of June 11, 2026.

Frequently Asked Questions

What is a VLA model?

A vision-language-action model is a neural network that maps camera images plus a natural-language instruction directly to robot actions — "pick up the cup" in, motor commands out. Architecturally it's usually a vision-language model (the kind that captions images) extended with an action head, so the robot inherits the VLM's world knowledge: it can handle objects and instructions it never saw in robot training because it saw them in internet pretraining.

Why are VLA models a big deal?

They broke robotics' scaling problem. Classical robotics hand-engineered each task; learned policies trained one task at a time. VLAs showed that one model, trained on diverse robot demonstrations on top of internet-scale pretraining, generalizes across tasks, objects, and even robot bodies — the same recipe that made language models general. That result is why "robot foundation model" companies raised billions before commercial deployment.

What are the most important VLA models?

The lineage: Google DeepMind's RT-2 (2023) proved the concept by bolting action output onto a VLM. OpenVLA (2024) open-sourced a competitive 7B version. Physical Intelligence's π series (π0 onward) became the flagship commercial line — and the company's researchers drove much of the original work. Figure's Helix runs humanoids with a fast-slow architecture, Google's Gemini Robotics brought frontier VLMs to manipulation, and Nvidia's GR00T line targets the same layer for the ecosystem.

What is the difference between a VLM and a VLA model?

A VLM (vision-language model) takes images plus text and outputs text — it can caption a photo or answer "what's on the table?". A VLA (vision-language-action model) takes the same images plus a text instruction but outputs robot actions — motor commands instead of words. Architecturally a VLA usually IS a VLM with an extra action head bolted on, fine-tuned on robot demonstrations. That lineage is the whole point: the VLA inherits the VLM's internet-scale world knowledge (what a "mug" is, what "left of the sink" means), which is why it can act on objects and instructions it never saw in robot training.

What is the difference between a VLA and a world model?

Direction. A VLA is a policy: situation in, action out. A world model is a simulator: situation plus action in, predicted next situation out. They're complementary — world models can train and evaluate VLAs in imagination instead of on expensive real hardware, which is why labs pursue both. The hottest 2025–2026 research thread is combining them: policies that plan inside a learned model of physics.

How do I keep up with VLA research?

It moves weekly. Our papers pipeline scores new robotics/AI research daily and summarizes the top one or two — VLA and embodied-AI work dominates the selection. The vision-language-action-models theme tracks the companies; the papers index has the research. Both are free and update automatically.

How does a VLA model actually turn an instruction into robot movement?

It runs in one forward pass. The model encodes the camera frames and the instruction text into a shared representation, the same way a vision-language model does to caption a photo, then an action head decodes that representation into motor commands rather than words. Newer designs such as Physical Intelligence's pi series use flow-matching action decoders to emit smooth, high-frequency continuous control instead of coarse discretized tokens, which is what lets the policy run on real hardware at usable speeds.

Are there any open-source VLA models I can use?

Yes. OpenVLA, released in 2024, is the reference open model: a 7B vision-language-action model trained on the multi-robot Open X-Embodiment dataset and competitive with closed systems, and most academic VLA work builds on it. Beyond that, the building blocks are open even where flagship weights are not — open vision-language-model backbones plus the public Open X-Embodiment demonstrations let teams fine-tune their own policies. Commercial flagship lines like the pi series and Helix are not fully open.

How much do experts on Teahose actually talk about VLA models?

A lot for a niche architecture term. Searching our summary corpus, vision-language-action models come up in 60 of the 1,150+ expert podcast, newsletter, and research conversations we have analyzed as of June 2026 — a share-of-voice signal that this is a load-bearing concept in physical AI, not a passing buzzword. We surface that discussion through the vision-language-action-models theme and the daily papers index.