Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/ARXIV PHYSICAL AI RESEARCH/Counterfactual Video Generation…
PAPR
// RESEARCH PAPER
ARXIV PHYSICAL AI RESEARCH

Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation

DATE September 29, 2026SOURCE ARXIV PHYSICAL AI RESEARCHPARTICIPANTS ZIHAN WANG, ANGJOO KANAZAWA, ET AL. (ARXIV PHYSICAL AI)ARXIV 2609.38172
In this episode
// SUMMARY

1. Key Themes

Data Scaling via Counterfactual Video Generation

PRISM uses video-to-video (V2V) generation to expand just four real-world videos into 256 diverse "counterfactual" interaction videos. This bypasses the costly and impractical task of filtering internet videos for clear, unoccluded full-body interactions. As the paper states, "Filtering Internet videos to find demonstrations that clearly show the person’s full body and how they interact with objects is therefore costly and impractical at scale" (Section 1). By generating variations of boxes, balls, bins, and barrels, the system creates behavior-level variations that are difficult to hard-code through geometry-level augmentation alone.

Contact-Anchored Real-to-Sim Pipeline

The core technical challenge is turning imperfect monocular video reconstructions into physically plausible robot trajectories. PRISM solves this by using human-object contact as a shared constraint across reconstruction, retargeting, and policy learning. "Our key insight is to use contact signal as a shared constraint across reconstruction, retargeting, and policy learning" (Section 1). During reconstruction, human motion guides the object trajectory while object contact corrects human pose errors. During retargeting, contact anchors specify where robot end-effectors should act, refining noisy reconstructions into physically plausible motions.

Zero-Shot Sim-to-Real Deployment on Humanoids

The framework culminates in a unified depth-based policy deployed on a real Unitree G1 humanoid without any real-world fine-tuning. The policy uses only onboard depth observations and joystick commands to pick up, carry, and drop diverse objects. "Using only onboard depth observations, our humanoid picks up, carries, and drops objects—including boxes, barrels, bins, and balls—across novel instances, sizes, and initial configurations" (Abstract). The system runs at 50 Hz, taking depth estimates from a head-mounted stereo camera.

Generalization to Unseen and Out-of-Domain Objects

The trained policy demonstrates robust generalization, not just to unseen instances of trained categories, but to entirely new object categories. In real-world tests, the policy successfully carried out-of-domain objects like tables, chairs, and backpacks with high success rates (Table 2). Furthermore, "Our policy succeeds zero-shot on a 35° ramp and a 0.43 m elevated support (Fig. 5), likely due to grasp-height overlap with tall training objects" (Section 5.3), showing emergent generalization to unseen terrain and elevations.

2. Contrarian Perspectives

Video Generation Models are Better Data Sources than Internet Videos

While the robotics industry often focuses on scraping internet video (like YouTube) for imitation learning data, this paper argues that internet videos are fundamentally misaligned with robot training needs. The paper notes that internet video "content and framing are shaped by human viewing preferences" making it "costly and impractical at scale" to find useful demonstrations (Section 1). Instead, using video generative models to synthesize "counterfactual" interactions from a few seed videos provides object-conditioned interaction strategies that adapt human behavior to the object's affordances, which is "difficult to hard-code through geometry-level augmentation alone" (Section 3.1).

Imperfect Reconstructions are Acceptable with Contact Constraints

Conventional real-to-sim pipelines strive for highly accurate 3D reconstructions, but PRISM demonstrates that perfect reconstruction is unnecessary. The authors explicitly state, "imperfect reconstruction is unavoidable in monocular real-to-sim stage, as no existing pipeline can recover physically plausible human–object motions" (Section 3.2). By relying on contact-anchored retargeting and rewards, the system can take "sufficiently close reconstructions" and optimize them into physically plausible robot demonstrations, bypassing the need for flawless 4D capture.

V2V Generation Outperforms Geometric Augmentation

A common approach to data augmentation in robotics is to apply geometric transformations (scaling, rotating, translating) to existing demonstrations. The paper challenges this by showing that V2V generation produces superior policies with fewer demonstrations. In their ablation, "V2V outperforms geometric augmentation with fewer demonstrations (80 vs. 103; Appendix F)" (Section 5.3). Specifically, V2V achieved 72.92% success on out-of-domain objects compared to 22.92% for geometric augmentation with more data (Table 9).

3. Companies Identified

Amazon FAR

Description: Amazon's Fundamental AI Research (FAR) team. Why relevant: The research was conducted primarily by Amazon FAR, indicating Amazon's significant investment and internal progress in humanoid robotics and embodied AI. The co-leads of the FAR team are authors on this paper. Quotes: "1Amazon FAR" (Author Affiliations); "† FAR Team Co-Leads" (Author Affiliations).

Unitree

Description: Manufacturer of the Unitree G1 humanoid robot. Why relevant: The real-world deployment and policy execution were demonstrated on the Unitree G1, a commercially available humanoid. This shows the framework's compatibility with leading hardware. Quotes: "We deploy our controller at 50 Hz on a 29-DoF Unitree G1" (Section 5.2).

NVIDIA

Description: Manufacturer of GPUs used for AI training. Why relevant: The training pipeline relies heavily on NVIDIA hardware, specifically 8 L40S GPUs, highlighting the compute requirements for scaling this type of real-to-sim-to-real pipeline. Quotes: "using 4096 environments per GPU on 8 NVIDIA L40S GPUs" (Section 4.2).

Intel

Description: Manufacturer of the RealSense D435i camera. Why relevant: The robot uses an Intel RealSense D435i for stereo image capture, which is then processed by Fast-FoundationStereo for depth estimation. This highlights the sensor stack used for zero-shot deployment. Quotes: "A head-mounted D435i camera captures stereo images at 30 Hz" (Section 5.2).

4. People Identified

Pieter Abbeel

Lab/Institution: Amazon FAR, UC Berkeley Why notable: A pioneer in robot learning and reinforcement learning. His involvement signals top-tier academic and industry convergence on humanoid loco-manipulation. Quotes: "Pieter Abbeel1,2†" (Author Affiliations).

Jitendra Malik

Lab/Institution: Amazon FAR, UC Berkeley Why notable: A leading figure in computer vision. His focus on visual imitation and real-to-sim pipelines underscores the importance of perception in physical AI. Quotes: "Jitendra Malik1,2†" (Author Affiliations).

C. Karen Liu

Lab/Institution: Amazon FAR, Stanford Why notable: Renowned for work in computer graphics and human motion synthesis, bringing expertise in physically plausible simulations to the robotics domain. Quotes: "C. Karen Liu1,4†" (Author Affiliations).

Guanya Shi

Lab/Institution: Amazon FAR, Carnegie Mellon University Why notable: Expert in control and learning for robotics, bridging the gap between dynamic locomotion and manipulation. Quotes: "Guanya Shi1,3†" (Author Affiliations).

Angjoo Kanazawa

Lab/Institution: Amazon FAR, UC Berkeley Why notable: Specialist in 3D human motion reconstruction from monocular video, crucial for the real-to-sim data acquisition phase of this research. Quotes: "Angjoo Kanazawa1,2†" (Author Affiliations).

5. Operating Insights

Leverage Contact as a Universal Constraint

For teams building manipulation pipelines from human video, explicit contact detection should be the anchor for the entire stack. Instead of treating reconstruction, retargeting, and policy learning as independent modules, PRISM uses contact points to couple human and object motion during reconstruction, guide end-effector placement during retargeting, and define rewards during RL. The paper notes, "human motion provides a strong prior for the object trajectory, while the object in turn constrains inaccurate human joint estimates" (Section 3.2). This approach salvages noisy monocular data that would otherwise fail in simulation.

Use Offboard Stereo Depth Estimation to Bridge Sim-to-Real

Deploying depth-based policies often suffers from a sim-to-real gap due to noisy or sparse depth sensors. PRISM mitigates this by running Fast-FoundationStereo on an external computer connected via Ethernet to process stereo images from a head-mounted camera. The authors found that "Fast-FoundationStereo substantially reduces the depth sim-to-real gap" (Section 5.2). For CTOs, this means investing in robust offboard perception processing can yield significant dividends in zero-shot deployment reliability over relying solely on native robot depth sensors.

Warm-Start Multi-Category Policies from Single-Category Training

When training a unified policy across multiple object categories, starting from scratch can be slow. PRISM demonstrates that training a box-only policy first, and then restoring only the actor weights to train on all categories, accelerates convergence. "Training from scratch succeeds; warm starting accelerates convergence and is used for our released checkpoint" (Section 4.2). This curriculum approach allows teams to rapidly iterate on core behaviors before scaling to diverse manipulation tasks.

6. Overlooked Insights

High Failure Rate in Trajectory Reconstruction

While the paper highlights the expansion from 4 videos to 256 generated clips, there is a significant drop-off during the real-to-sim phase. Of the 256 generated videos, only 137 yield feasible robot-object trajectories. "The remaining sequences fail the constrained solver, mainly because of severe collisions. In particular, some generated objects are too large for the Unitree G1 humanoid to carry without object–body collisions" (Appendix C). This indicates that V2V models still generate physically implausible scenarios for specific robot morphologies, and automated filtering or physics-based validation of generated videos remains a bottleneck.

Payload Sensitivity for Out-of-Domain Objects

Although the policy generalizes to out-of-domain objects, its success rate degrades significantly with heavier payloads. While in-domain success only drops from 98.75% (0.1-1.0kg) to 91.25% (4.0-5.0kg), out-of-domain success drops from 77.08% to 64.58% for the same weight increase (Table 8). This suggests that while the policy can grasp novel geometries, it lacks robustness to the dynamic shifts caused by heavy, unmodeled mass distributions, limiting immediate deployment in industrial settings with heavy parts.