Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/ARXIV PHYSICAL AI RESEARCH/Long-WAM: Scaling the Context of…
PAPR
// RESEARCH PAPER
ARXIV PHYSICAL AI RESEARCH

Long-WAM: Scaling the Context of World-Action Models

DATE October 7, 2026SOURCE ARXIV PHYSICAL AI RESEARCHPARTICIPANTS WEI HUANG, YUKANG CHEN, ET AL. (ARXIV PHYSICAL AI)ARXIV 2610.10528
// SUMMARY

1. Key Themes

Autoregressive Video Pretraining Unlocks the Value of Long Context

The paper's central finding is that simply giving a robot policy more visual history does nothing unless the underlying video foundation was trained autoregressively (predicting the future from the past, frame by frame). When the authors compared three video foundation initializations—bidirectional (Wan2.2), general AR (LongLive-2.0), and robot-domain AR (LongLive2.0-Robot)—only the AR variants gained from longer history. On RoboCasa GR-1, the robot-domain AR model's lead over the bidirectional initialization grew from 3.3 points with no history to 17.1 points at 19.2 seconds of history. The bidirectional variant actually went from 64.1% at 9.6 seconds back down to 61.6% at 19.2 seconds—more history hurt it. As the paper states: "access to history is not the same as using it: longer histories pay off far more when the video foundation is pretrained autoregressively (AR)" (Abstract). The practical implication: companies building world-action models on top of bidirectional video generators (the current dominant approach) may be hitting a ceiling that AR pretraining could break through.

Real-Time Edge Deployment with Full Future Prediction at 107 ms

The paper demonstrates that you can run a full world-action model—including future video latent prediction—on edge hardware fast enough for closed-loop control. On an RTX 5090, each action chunk takes 107.4 ms including the complete observation VAE computation and future-video latent prediction (Section 4.2, Table 8). This is a 3.3× speedup over unoptimized BF16 execution (356.0 ms). The system also deploys on DGX Spark (328.2 ms) and Jetson AGX Thor (378.7 ms). The key engineering innovations are streaming VAE encoding (processing incoming frames as they arrive rather than waiting for the full window), asynchronous execution without action blending, NVFP4 quantization for the video expert while keeping action compute in BF16, and device-specific kernel tuning. This matters because it proves that "imagine-then-act" pipelines—where the robot first predicts what will happen visually, then generates actions conditioned on that prediction—can run at usable speeds on real robot hardware.

Dramatic Superiority on Dynamic Manipulation Tasks

On a Unitree G1 humanoid, Long-WAM achieves 90–100% grasping success across conveyor speeds from 3.0 to 7.5 cm/s, while Pi0.5 drops from 15% to 0% and Fast-WAM drops from 75% to 0% (Section 6.1, Figure 7). On dynamic cup stacking (intercepting a moving cup and nesting it inside another), Long-WAM succeeds in 19 of 20 trials (95%), whereas "neither π0.5 nor Fast-WAM succeeds once" (Abstract). This is not a marginal improvement—it is a categorical difference in capability. The paper attributes this to the model's ability to use recent visual history to anticipate object motion and time its interception, which is exactly what the IDM (inverse dynamics modeling) architecture provides: first predict how the scene will evolve, then generate actions conditioned on that prediction.

Strong Execution Policy Amplifies High-Level Planning

When paired with GPT-6 Astra as a high-level planner on RoboCasa365 composite tasks, Long-WAM's overall success jumped from 31.4% to 54.4%, while Pi0.5 paired with the same planner only went from 16.9% to 30.5% (Section 6, Table 5). The planner augmentation gain for Long-WAM was +23.0 points versus +13.6 for Pi0.5. On unseen skill compositions, Long-WAM + planner achieved 35.0% versus 19.0% for Pi0.5 + planner. The paper's interpretation: "stronger execution lets planning pay off more" (Section 1). This suggests that the bottleneck in hierarchical robot systems may be execution quality, not reasoning capability—and that investing in a strong low-level controller may yield better returns than investing in smarter planners.


2. Contrarian Perspectives

Bidirectional Video Pretraining Is the Wrong Foundation for Robot Control

Most current world-action models (DreamZero, LingBot-VA, Motus) adapt bidirectionally pretrained video generators for causal control. The paper directly challenges this: "DreamZero [52] and LingBot-VA [26] adapt bidirectionally pretrained video generators to causal video–action prediction" and shows that this approach cannot exploit longer history (Section 1, Figure 6). The bidirectional initialization showed "no net gain" from 0 to 19.2 seconds on GR-1 (Abstract). The reasoning is that AR pretraining learns causal temporal factorization—predicting the future from the past—during the video stage, which is exactly the structure WAM adaptation needs. Bidirectional models learn to fill in missing frames given both past and future context, which does not transfer to the strictly causal prediction needed for real-time control. Companies building on bidirectional video foundations may be paying a hidden tax: their models can access history but never learned to predict from it.

Future Video Prediction Is Worth the Compute—Contradicting Fast-WAM

Fast-WAM (a prior work by Yuan et al., 2026) explicitly questioned whether world-action models need test-time future imagination. Long-WAM argues the opposite: future video prediction is essential for dynamic tasks. The IDM variant (which predicts future video latents before generating actions) achieves 99.5% on LIBERO-Long versus 94.5% for the no-future-prediction variant (Table 1), and the real-world dynamic manipulation results show an even starker difference—95% vs 0% on dynamic cup stacking (Section 6.1). The paper also shows that IDM's sequential cost is manageable: 107.4 ms on RTX 5090, which is actually 2.3× faster than Fast-WAM's 244.1 ms despite "additionally predicting future video frames" (Section 6.2, Table 6). The implication: dropping future prediction to save compute is a false economy for dynamic manipulation.

Pure Asynchronous Execution Without Blending or Prefix Guidance Is Sufficient

The robotics control literature has developed elaborate mechanisms—inference-time guidance, prefix-conditioned training, denoising-time blending—to ensure smooth transitions between action chunks in asynchronous execution. Long-WAM requires none of these: "Long-WAM requires neither in our experiments" (Section 4.1). On RoboTwin 2.0, Long-WAM (IDM) loses only 0.2 percentage points going from synchronous to asynchronous execution (94.4% to 94.2%), while Fast-WAM loses 15.4 points and LingBot-VA loses 45.5 points (Table 7). The paper attributes this to "temporally continuous LongLive2.0-Robot forecasts" producing consecutive chunks that naturally agree in their overlap. This suggests that the smoothness problem in action chunking may be a symptom of weak predictive models, not an inherent engineering challenge requiring complex mitigation.


3. Companies Identified

NVIDIA, Hardware and research powerhouse, Primary affiliation of most authors including senior researchers; deployment targets are all NVIDIA hardware (RTX 5090, DGX Spark, Jetson AGX Thor). The paper essentially serves as a showcase for NVIDIA's full-stack Physical AI strategy—from Blackwell architecture (NVFP4 quantization) to CUDA Graphs to edge deployment. "Together, shared optimizations and device-specific tuning yield 3.2–4.1× total speedups over BF16 eager execution" (Section 4.2).

Physical Intelligence (Pi / Pi0.5), Robotics foundation model company, Creator of π0 and π0.5, used as a primary baseline throughout. Pi0.5 fails completely on dynamic cup stacking (0/20 trials) and achieves only 15% grasping success at 3 cm/s conveyor speed, dropping to 0% at higher speeds. When paired with GPT-6 Astra as a planner, Pi0.5 + planner achieves 30.5% overall on RoboCasa365 versus 54.4% for Long-WAM + planner (Table 5). The paper positions Pi0.5 as a strong generalist that nonetheless lacks the dynamic manipulation capability that predictive history enables.

Unitree, Humanoid robot manufacturer, Manufacturer of the G1 humanoid used for real-world dynamic manipulation evaluation. The paper demonstrates 90–100% grasping success and 95% dynamic stacking success on the G1 (Section 6.1, Figure 7). This validates the G1 as a capable platform for dynamic manipulation research.

AgiBot, Robotics data and platform company, Creator of AgiBot World dataset, which contributes 4.8% of the pretraining corpus (110,331 samples across two subsets, Table 9). The dataset is used for the robot-domain AR video pretraining stage.

Physical Intelligence (π0.5) and the broader VLA ecosystem, The paper benchmarks against a wide field including OpenVLA, GR00T-N1 (NVIDIA's own), UniVLA, X-VLA, InternVLA-M1, StarVLA, and Qwen-RobotManip. Long-WAM outperforms all of these on LIBERO (99.5% average) and most on RoboTwin 2.0 (94.4% average). The competitive landscape is crowded, but Long-WAM's advantage is specifically in dynamic and long-horizon tasks where visual history matters most.


4. People Identified

Linxi "Jim" Fan, NVIDIA, Senior research scientist at NVIDIA and well-known figure in Physical AI (co-created MineDojo, Eureka, GR00T). Co-author on this paper, contributing to the world-action modeling framework. His involvement signals NVIDIA's strategic commitment to world-model-based robot control.

Song Han, MIT, Professor at MIT renowned for model compression, quantization, and efficient deep learning. Co-author and likely driving force behind the NVFP4 quantization, CUDA Graph optimization, and device-specific kernel tuning that enable real-time edge deployment. His expertise bridges the gap between large video models and edge compute constraints.

Xiaojuan Qi, HKU (University of Hong Kong), Associate Professor at HKU working on computer vision and robotics. Co-author and corresponding contributor to the world-action modeling framework. HKU is emerging as a significant contributor to Physical AI research.

Wei Huang and Bohan Zhang, NVIDIA and MIT (equal contribution), Lead authors of the paper. Huang is affiliated with NVIDIA and Zhang with MIT. Their work spans both the model architecture (AR pretraining, causal-to-causal adaptation) and the deployment system (streaming VAE, asynchronous execution).

Yukang Chen, NVIDIA, Co-author who also led LongLive-2.0 (the AR video generation infrastructure that Long-WAM builds upon). The lineage from LongLive-2.0 to Long-WAM represents NVIDIA's sustained investment in autoregressive video generation for robotics.


5. Operating Insights

Choose Your Video Foundation Based on Whether You Need Temporal Reasoning

If your robot operates in static environments where the current observation is sufficient, the choice of video pretraining (AR vs. bidirectional) matters less. But if your use case involves moving objects, multi-step interactions, or any task where recent history informs the next action, AR pretraining is essential. The data is stark: on RoboCasa GR-1, the AR advantage grows from 3.3 points with no history to 17.1 points at 19.2 seconds (Section 5.3, Figure 6). The bidirectional model actually gets worse with more history. For a CTO choosing a video backbone for a world-action model, this is a foundational architectural decision that cannot be patched later. The paper also notes that "robot-domain AR pretraining further raises peak success on both LIBERO-Long and GR-1" (Abstract)—meaning that pretraining on robot-specific video, not just general video, provides additional gains.

Context Length Should Be Task-Dependent, Not Maximized

The paper shows that different tasks have different optimal history windows: LIBERO-Long saturates at 2.4 seconds (94.5% → 99.5%), while RoboCasa GR-1 keeps improving up to 19.2 seconds (63.3% → 78.7%) (Section 5.3, Figure 6). Going to 38.4 seconds actually hurts performance (drops to 75.2%), partly because 80.4% of sampled history frames are padding at that length. The latency cost is also significant: 8× more history raises RTX 5090 latency 3.2× (107.4 to 341.0 ms; Appendix G, Table 12). The practical takeaway: don't default to the longest context your hardware can handle. Profile your task's memory needs, and consider adaptive context allocation—the paper explicitly motivates "a single policy that adjusts its window at test time" as the natural next step (Appendix A).

The Execution Policy Is the Bottleneck for Hierarchical Systems, Not the Planner

When evaluating Long-WAM + GPT-6 Astra on RoboCasa365, the planner augmentation yielded +23.0 points for Long-WAM but only +13.6 for Pi0.5 (Section 6, Table 5). On unseen compositions, Long-WAM + planner achieved 35.0% versus 19.0% for Pi0.5 + planner. The paper's interpretation: "Planner augmentation yields a larger gain for Long-WAM (+23.0 points) than for π0.5 (+13.6), suggesting that stronger execution lets planning pay off more" (Section 1). For companies building hierarchical robot systems, this means investing in a stronger low-level executor may deliver better ROI than investing in a smarter planner. The planner can only compose skills that the executor can reliably perform.


6. Overlooked Insights

Action Compute Stays in BF16 While Video Goes to 4-Bit—A Deliberate Precision Choice

The paper makes a specific precision tradeoff that is easy to miss: video expert linear layers use NVFP4 (4-bit weights and activations) while "action compute and KV storage remain BF16" (Section 4.2). The rationale: "Action quantization offers limited latency savings and may compromise control precision, so we retain BF16." This is a critical engineering decision for anyone deploying quantized robot policies. The video prediction can tolerate aggressive quantization because it serves as a conditioning signal, but the action output directly controls the robot and small numerical errors could compound. Companies deploying VLA or WAM models on edge hardware should consider this asymmetric quantization strategy rather than uniformly quantizing the entire model.

The 38.4-Second Context Decline Is a Data Problem, Not a Model Limitation

When context was extended to 38.4 seconds, success dropped from 78.7% to 75.2% on GR-1 (Section 5.3). The paper notes that this window is "three times the average training trajectory (12.1 seconds), and 80.4% of its sampled history frames are padding" and hypothesizes "the decline reflects limited history coverage rather than an intrinsic memory limit" (Section 5.3). This is a significant insight for the industry: the context scaling ceiling may be set by dataset quality (trajectory length and density), not model architecture. Companies investing in longer-horizon robot data collection—specifically, longer continuous demonstrations—may be directly enabling longer effective context windows for their policies. The model architecture is ready; the data is the bottleneck.