Vision-Language-Action Models
Companies developing vision-language models that are extended with action heads or policies to directly control physical robots and autonomous systems.
CAPITAL FIGURES ARE MEDIA-EXTRACTED ESTIMATES, NOT VERIFIED FILINGS.
EXTRACTED FROM 25+ PODCASTS & VC NEWSLETTERS · MEDIA-REPORTED FIGURES, NOT VERIFIED FILINGS
Gemini Robotics 2 resets the frontier VLA benchmark
Google DeepMind's Gemini Robotics 2 (signal [1]) and the concurrent release of Gemini 3.1 Pro (signal [41]) mark a step-change in foundation-model-to-robot-policy transfer, with AGI countdown sites revising probability upward to 98% on its announcement (signal [4]). Gemini 3.1 Flash already leads competing VLM architectures on seed-point prediction at 71.87% average success versus Qwen 3.5 Flash at 54.42% and GPT 5.4 at 42.61% (signal [21]), establishing Google DeepMind as the performance leader for real-world manipulation tasks. The release validates the thesis that large vision-language backbones can be adapted end-to-end into robot controllers—a design principle Google DeepMind has championed—and sets a new competitive baseline that every other VLA lab must now respond to.
π0.5 was formally published as an arXiv product release (signal [34]), with research confirming that representations from its pretrained backbone enable rapid improvement throughout training (signal [31]), and it is now used as the basis for derivative architectures like XS-VLA's action-generation approach (signal [43]). Temporal GRPO methods are already being benchmarked against π0, with Temporal GRPO scoring 75.8% versus π0's 49.2% on RoboTwin 2.0 (signal [3]), signaling that the open research community is actively building on and surpassing it. Physical Intelligence's dual role as both a product company and a research platform provider mirrors Hugging Face's dynamic in language AI.
Why it matters · A widely adopted open backbone accelerates the entire ecosystem but compresses the window for Physical Intelligence to monetize proprietary performance advantages.
NVIDIA appears as a co-investor in multiple large rounds—including a $2B growth round (signal [14]) and a separate $1.1B round alongside AMD Ventures (signal [7]) and another $1.1B round with General Catalyst, YC, and Temasek (signal [24])—while simultaneously open-sourcing Cosmos (including training frameworks, synthetic data, and model weights, signal [17]) and releasing Cosmos 3, an omni-modal World Foundation Model combining video, audio, language, and action signals (signal [19]). NVIDIA's GR00T N1 is cited as the baseline generalist robot foundation model, and Temporal GRPO post-training methods are explicitly described as applicable to GR00T (signal [11]). Hardware benchmarking across the arXiv Physical AI corpus uniformly relies on NVIDIA silicon—RTX PRO 6000, Jetson Thor, RTX 3090/4090/5090 (signal [9])—cementing compute lock-in alongside the software stack.
Why it matters · NVIDIA is positioning to capture value at every layer of the VLA stack—silicon, simulation, foundation models, and venture capital—making it the unavoidable platform for the physical AI era.
SmolVLA (signal [44]) and SmolVLM2-0.25B (signal [42])—both sub-3B parameter architectures—demonstrate that capable VLA policies can be distilled into models small enough for edge deployment, directly contesting the industry assumption that scaling alone yields deployable robots (signal [45]). XCoT-VLA introduces executable chain-of-thought reasoning for autonomous driving VLAs (signal [39]), while research explicitly flags that verbose natural-language CoT is poorly suited to real-time control due to decoding cost and optimization difficulty (signal [40]). These architectural innovations point toward a bifurcated market: massive frontier VLAs for training and a new tier of compact, deployment-ready action models.
Why it matters · Operators and OEMs deploying robots at scale will favor efficient edge-deployable VLAs, opening a new product category distinct from frontier research models.
Weekly capital data shows deal sizes spiking dramatically—$13.8B in the week of July 13 and $11.5B the week of July 27—against a 90-day total of $24.6B across 29 deals, with 'unknown' stage rounds ($27.7B across 31 deals) and Series A ($12.5B) dominating the stage mix, suggesting both late-stage growth financings and early commercialization bets are occurring simultaneously. The $10B Berkshire Hathaway Series Growth round (signal [49]) and the $2B growth round backed by Blackstone, Jane Street, Coatue, and NVIDIA (signal [14]) reflect non-traditional deep-pocketed investors treating VLA infrastructure as durable productive assets rather than venture speculation. Series C rounds at $9.7B collectively further confirm that multiple companies are crossing from research to revenue.
Why it matters · When insurers, asset managers, and sovereign funds lead VLA rounds, the asset class has crossed the threshold from speculative to infrastructure-grade—compressing the timeline for incumbents to establish defensible positions.