AI Robot Manipulation
Companies building AI-native systems that enable robots to perceive, grasp, and dexterously manipulate objects in unstructured real-world environments.
CAPITAL FIGURES ARE MEDIA-EXTRACTED ESTIMATES, NOT VERIFIED FILINGS.
EXTRACTED FROM 25+ PODCASTS & VC NEWSLETTERS · MEDIA-REPORTED FIGURES, NOT VERIFIED FILINGS
VLA model wars intensify: π0.5, GR00T, and Gemini Robotics 2 collide
The vision-language-action model race has entered a decisive phase, with Physical Intelligence's π0.5 (built on flow-matching policy architecture), NVIDIA's GR00T-N1.5-3B, and Google DeepMind's Gemini Robotics 2 all shipping within the same week. Benchmarks are now the battlefield: Temporal GRPO methods score 75.8% on RoboTwin 2.0 versus Physical Intelligence's π0 at 49.2%, while RoboBRIDGE improves average success rates from 3.7% to 7.5% on RoboCasa across multiple VLA backbones. The open-weight release of π0.5 and NVIDIA's open-sourcing of Cosmos — including training frameworks, synthetic data, and model weights — are commoditizing the base layer and forcing differentiation up the stack. Critically, a counter-thesis is gaining traction: researchers now argue that scaling VLAs alone won't yield deployable robots, and that orchestration frameworks are the missing layer.
NVIDIA is no longer just a chip supplier to robotics — it is simultaneously the benchmark hardware (RTX PRO 6000, Jetson Thor used for all evaluations), the simulation platform (Isaac Sim), the foundation model provider (GR00T N1, Cosmos 3 World Foundation Model), and increasingly a capital-markets actor raising $500B via Goldman Sachs, BlackRock, and others. NVIDIA's VP Liu Mingyu explicitly frames physical AI as a market-creation exercise analogous to CUDA's role in the software AI boom. With NVIDIA also appearing as a lead investor in a $2B growth round (alongside Blackstone, Jane Street, and Coatue) and a $1.1B Series B, the company is placing capital bets across the stack it also supplies.
Why it matters · NVIDIA's vertical integration across silicon, simulation, foundation models, and financing creates a gravitational pull that makes it the de facto platform tax on every robotics startup — operators must decide whether to build on or around it.
Capital deployment into the physical AI stack has become episodic and enormous: weekly deal values swung from $504M (week of July 20) to $17.8B (week of July 6) and $12.3B (July 13), with $23.7B deployed across just 30 deals in the last 28 days. The stage mix reveals a bifurcated market — 28 'unknown' rounds totaling $33.3B sit alongside 26 Series A rounds at $13.1B, suggesting both early formation activity and large structured/growth financings are occurring simultaneously. GPU securitization mechanisms and $500B infrastructure debt facilities backed by Apollo, BlackRock, Blackstone, and KKR are effectively industrializing the capital formation process for physical AI compute.
Why it matters · The financialization of AI compute infrastructure means robotics companies with NVIDIA-backed hardware can now access sovereign-wealth-grade capital, dramatically lowering the cost of scaling hardware fleets.
NVIDIA Isaac Sim has become the default simulation environment cited in peer-reviewed robotics benchmarks, and Cosmos 3 — progressing from version 1 to version 3 in under 18 months — is positioned as an omni-modal World Foundation Model combining video, audio, language, and action signals. Sunday Robotics is building humanoid robots with a dedicated teleoperation training data pipeline, while the Columbia RoboPIL Lab and UC Berkeley continue publishing on diffusion-based policies and deformable material handling that feed into commercial sim-to-real transfer.
Why it matters · As Isaac Sim and Cosmos commoditize simulation infrastructure, the competitive advantage moves to proprietary real-world demonstration data — the resource that synthetic pipelines cannot fully replicate.
A growing body of arXiv research directly challenges the assumption that scaling VLA models alone will produce deployable robots, arguing instead that orchestration frameworks — systems that coordinate perception, planning, failure diagnosis, and action primitives — are the missing architectural layer. Evaluations show even Gemini-3 Flash exhibits only marginal improvement as a planner/monitor backbone, exposing the gap between raw model capability and robust task execution. This is opening space for companies like Covariant and Sereact, which focus on intelligent manipulation systems and warehouse orchestration rather than frontier model development.
Why it matters · Startups that can ship reliable orchestration layers on top of open-weight VLAs (π0.5, GR00T, OpenVLA) will capture enterprise deployment budgets that pure model labs cannot address.