Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/ARXIV PHYSICAL AI RESEARCH/InternW0-Δ: A World Action Model…
PAPR
// RESEARCH PAPER
ARXIV PHYSICAL AI RESEARCH

InternW0-Δ: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data

DATE September 25, 2026SOURCE ARXIV PHYSICAL AI RESEARCHPARTICIPANTS XINGYU MIAO, CHUNHUA SHEN, ET AL. (ARXIV PHYSICAL AI)ARXIV 2609.31394
// SUMMARY

1. Key Themes

Decoupling Future Video Generation from Action Inference

The paper introduces "Causal Imprint," a mechanism that learns future-relevant scene changes during training but does not require generating future video at inference time. Many World Action Models (WAMs) waste compute by generating future video frames to condition action prediction. By forcing the model to learn predictive representations from observed context using future outcomes only as a training target, they drastically reduce inference latency and compute costs. As stated in the paper, "Causal Imprint learns future-relevant scene changes from training-only future supervision and makes these representations available to the action expert without requiring future-video rollout at inference" (Abstract).

Scaling Open-Source Embodied Data to 20,000+ Hours

They curated a massive, heterogeneous dataset of over 20K hours of processed training data, combining robot demonstrations, UMI (Universal Manipulation Interface) data, egocentric human demonstrations, and Ego2Robot data. This scale of open-source embodied data is unprecedented. "The resulting corpus contains over 20K hours of processed training data, to our knowledge the largest open-source corpus of its kind" (Abstract). This provides a massive foundation for pretraining generalist robot policies.

Unified Canonical Representation for Cross-Embodiment Learning

To make use of such diverse data (ranging from single-arm robots to bimanual humanoids to human hand videos), they map everything into a single 80-dimensional canonical state-action space. This prevents the model from getting confused by different coordinate conventions or action dimensions across datasets. "We therefore convert all trajectories into a canonical state–action representation with fixed semantic slots" (Section 4.1). This allows a single pretrained model to be adapted to entirely different robot hardware.

High-Frequency Real-Robot Deployment via Inference Optimization

Real-world robotics requires fast control loops. The team optimized their inference pipeline to achieve a 152.8 ms round-trip latency on a single consumer-grade NVIDIA RTX 5090 GPU, a 5.11x speedup over standard runtime, enabling 30 Hz control rates. "The optimized runtime achieves an average controller-observed round-trip latency of 152.8 ms on a single NVIDIA RTX 5090 GPU, corresponding to a 5.11× speedup over the standard runtime" (Section 1). This proves the architecture is viable for edge deployment, not just cloud-based inference.

2. Contrarian Perspectives

Future Video Generation is Not Necessary for Effective Control

Many World Action Models (WAMs) require sampling future video to condition action generation, under the assumption that a model must "imagine" the future to act on it. This paper argues that explicit future generation at test time is computationally wasteful and unnecessary. By using "Causal Imprint" to distill future-relevant changes into current representations, they separate future supervision from inference. "This allows InternW0-∆ to directly predict actions at inference without sampling future videos or invoking the distillation branch" (Section 1). This challenges the prevailing WAM architecture paradigm by prioritizing action-relevant latent features over explicit video generation.

Egocentric Human Data Can Be Reliably Converted to Robot Training Data

Converting human hand videos to robot trajectories is often considered too noisy or unreliable due to kinematic mismatches and differing morphologies. This paper demonstrates a rigorous pipeline (Ego2Robot) that couples base search with inverse kinematics and visual alignment, yielding 5,633 robot-hours from human video. "Kinematic alignment couples base selection with inverse kinematics (IK)... Visual alignment uses SAM3 to segment human regions and ProPainter to remove them... The resulting videos are paired with the corresponding smoothed end-effector targets and gripper commands" (Section 4.3.2). This suggests that the robotics community is vastly underutilizing available human video data.

Heavy Data Filtering Yields Better Models Than Raw Scale

Instead of just dumping raw data into the model and hoping scale solves quality issues, they apply aggressive, multi-stage filtering (signal anomaly, visual quality, action magnitude, instruction consistency) which reduced robot data from 13,867 hours to 11,302 hours. "We apply signal anomaly and consistency filtering, static boundary trimming, visual quality filtering, and an action magnitude guard... We additionally perform automated checks of instruction correctness and video–instruction consistency" (Section 4.2.2). This argues that data curation and quality control are more important than simply maximizing raw hours of training data.

3. Companies Identified

Shanghai Artificial Intelligence Laboratory

Description: The institution behind the paper. Why relevant: They are open-sourcing the entire stack (code, weights, data pipeline) which could shift the competitive landscape for foundation models in robotics by lowering the barrier to entry. "We will open source the training code, model weights, infrastructure, data-processing pipeline, and processed data where licenses permit" (Abstract).

NVIDIA

Description: Hardware provider. Why relevant: The model is trained on 256 NVIDIA A800 GPUs and deployed on a single NVIDIA RTX 5090 GPU. Their hardware is the explicit deployment target for this edge-optimized model. "We pre-train our model on 256 NVIDIA A800 GPUs... Action inference runs on one RTX 5090 GPU" (Section 5.1 & 6.2).

AgiBot, Fourier, Galaxea, RealSource

Description: Robot hardware and data companies whose datasets are used for pretraining. Why relevant: Their data contributes to the 20K hour corpus, showing which hardware providers are generating usable manipulation data at scale. "AgiBotWorld is a large-scale real-world manipulation dataset... ActionNet focuses on dexterous bimanual manipulation using humanoid robots... Galaxea Open-World Dataset contains more than 500 hours..." (Section 4.2.1).

4. People Identified

Xingyu Miao et al. (48 total authors)

Lab/Institution: Shanghai AI Laboratory (Physical Intelligence Team) Why notable: This is a massive team effort from Shanghai AI Lab's Physical Intelligence team, indicating a strong institutional push into Physical AI foundation models. The scale of the author list (48 total) reflects the infrastructure-heavy nature of modern robotics foundation models. "Physical Intelligence Team, Shanghai Artificial Intelligence Laboratory" (Title page).

5. Operating Insights

Design for Inference Speed from the Ground Up

CTOs should note that the architecture explicitly avoids test-time video generation. By caching visual context and only recomputing action tokens during denoising, they achieve 30Hz control on a single edge GPU. "Inference consists of a single video-expert prefill followed by iterative action-only denoising... Only the action stream is recomputed during denoising" (Section 3.8). This directed information flow is key to deploying large transformer-based policies on real robots without unacceptable latency.

Canonical Action Spaces Enable Plug-and-Play Post-Training

By mapping all data to an 80-dimensional canonical space, the same pretrained checkpoint can be rapidly adapted to entirely different robot setups (grippers, dexterous hands) with minimal post-training data. "We post-train InternW0-∆ on 4 robot setups, including two gripper-based ones and two dexterous-hand ones... all post-training share the same settings and start from the same pretrained checkpoint" (Section 5.3). This means a company can invest in a single foundation model and adapt it to multiple product lines.

6. Overlooked Insights

Aggressive Trajectory Trimming and Speed Alignment

The team artificially slows down human egocentric data by a factor of two to match robot motion speeds, and trims static boundaries aggressively. This simple alignment of temporal dynamics across data sources is crucial for stable policy learning and is often overlooked in favor of more complex architectural tweaks. "To better match the motion speed of the robot data, we slow down EgoVerse and EgoDex trajectories by a factor of two" (Section 4.3.2).

Prefix-Conditioned Training for Asynchronous Execution

To handle the reality of inference delay, they train the model to predict action continuations given a known prefix of already-executed actions. This avoids complex inference-time guidance or action blending, baking the reality of latency directly into the training objective. "Following training-time RTC... we simulate inference delay by sampling a committed prefix length... This trains the policy to predict an action continuation conditioned on a known prefix" (Section 5.3).