World Motion Models: Flexible Sequence Modeling of SE(3) Trajectories
- 01One Model, Six Tasks: Unified SE(3) Representation Eliminates Task-Specific Architectures
- 02Sparse SE(3) Trajectories as a Universal Motion Language
- 03Per-Token Noise Levels Enable Any-to-Any Conditioning
- 04Cross-Embodiment Retargeting at 87 FPS with Physics-Plausible Output
- 05MPC Optimization Improves Policy Success and Speed
1. Key Themes
One Model, Six Tasks: Unified SE(3) Representation Eliminates Task-Specific Architectures
The core contribution is a single generative model architecture that handles policy learning, model-predictive control, 3D scene future prediction, human-object interaction generation, cross-embodiment retargeting, and motion planning — all by changing which tokens are "known" vs. "unknown" at inference time. The paper demonstrates this across six diverse benchmarks. As stated in the introduction: "Tasks such as future prediction, motion infilling, model-predictive control, inverse kinematics, cross-embodiment retargeting, and policy learning all reduce to the application of different masks over the same network." This is significant because today's robotics stacks typically require separate models for perception, planning, control, and simulation — WMM collapses all of these into one set of weights.
Sparse SE(3) Trajectories as a Universal Motion Language
The paper's foundational insight is that complex 4D world dynamics can be compactly represented as sparse sets of rigid SE(3) pose trajectories — replacing thousands of mesh vertices with tens of joint poses. From Section 2.1: "the dense flow of thousands of mesh vertices on a human or robot collapses into tens of joint trajectories; cameras and rigid objects track as single moving frames; and even non-rigid scene dynamics can be represented as interpolations of a small SE(3) motion basis." This representation unifies humans, robots, hands, objects, cameras, and scene dynamics into one shared token space without category-specific structure.
Per-Token Noise Levels Enable Any-to-Any Conditioning
The technical mechanism that makes the unified model work is assigning each token (one entity at one time step) its own independent noise level during flow-matching training. From Section 2.2: "each token (one entity at one time step) is assigned its own scalar level, sampled i.i.d." This means at inference, you can set noise=0 on observed tokens and noise=1 on target tokens, and the same network performs prediction, policy, or planning depending on the mask. The ablation in Table 6 shows that removing per-token noise (using synchronized noise) yields zero successful rollouts on the Language Table benchmark — the mechanism is not optional, it is essential.
Cross-Embodiment Retargeting at 87 FPS with Physics-Plausible Output
The paper demonstrates human-to-humanoid retargeting from monocular video to a Unitree G1, achieving 86.97 FPS inference while maintaining higher sequence success rates (0.8690) than both the real-time baseline GMR (0.7945) and the offline optimization method PHUMA (0.8255). From Table 5, WMM also handles noisy input (σ=0.1 per-frame) and sparse keyframe infilling (3 FPS) with a single checkpoint, while PHUMA's quality degrades significantly under sparse keyframes (success drops from 0.8255 to 0.6535).
MPC Optimization Improves Policy Success and Speed
When WMM is used as a differentiable world model for model-predictive control rather than as a plain policy, success rates improve and rollouts complete faster. From Table 2: on B2B tasks, MPC achieves 100% success vs. 90% for plain policy, and on B2AL, 80% vs. 76%. Minimum steps to success drop from 26 to 16 (B2B) and from 14 to 2 (B2AL), meaning the MPC-optimized policy reaches goals far more efficiently.
2. Contrarian Perspectives
Video-Based World Models Are the Wrong Abstraction for Robotics
Most leading world model efforts (NVIDIA Cosmos, V-JEPA 2, DreamDojo, World-VLA-Loop) operate in pixel/video space, implicitly capturing 3D dynamics through color changes on a 2D screen. WMM explicitly argues this is fundamentally entangled: "It is also orthogonal to video-based world modeling, which captures dynamics implicitly through color changes on a screen, entangling the underlying 3D motion with appearance, lighting, and viewpoint changes" (Section 1). The paper's position is that SE(3) pose trajectories are a more principled and compact primitive — a 30-frame, 400-trajectory scene is only 12K tokens, far more tractable than video frames. Companies betting on video world models for robot control may be solving a harder problem than necessary.
Morphology-Specific Models Are a Dead End
The robotics industry is largely organized around embodiment-specific models — separate architectures for arms, humanoids, hands, and mobile bases. WMM argues this fragmentation is unnecessary: "We further impose no category-specific structure: every entity is treated equally as a pose proxy, generalizing class-bound representations (SMPL, MANO, URDFs) to any unstructured scene" (Section 1). The paper demonstrates motion planning across 8+ different manipulators (UR3, iiwa14, iiwa7, kinova, panda, UR5, UR10, widowx) within a single model (Figure 8), suggesting that the industry's tendency to build embodiment-specific stacks may be leaving significant generalization on the table.
Non-Causal Inference Is as Important as Causal Prediction
Standard world models and policies are strictly causal — predict the future from the past. WMM argues that real-world robotics requires non-causal inference just as often: inverse dynamics (what action achieves this goal state), motion infilling (plan between waypoints), and cross-embodiment retargeting (infer correspondences between trajectories). From Section 1: "humans can perform non-causal inference as well, such as determining the actions required to achieve a specific goal (inverse dynamics) or planning intermediate motions between waypoints." The paper's context token mechanism allows goal states arbitrarily far in the future to condition current planning, which standard causal architectures cannot do without ad hoc workarounds.
3. Companies Identified
Unitree Robotics
- Description: Manufacturer of the G1 humanoid robot
- Why relevant: The G1 is used as the target embodiment for human-to-humanoid retargeting experiments (Section 3.4). The paper demonstrates real-world monocular video retargeting to the G1, showing that WMM can serve as a motion translation layer between human demonstrations and humanoid robots — a critical capability for companies using teleoperation or human motion capture to train humanoid policies.
- Quote: "We construct a benchmark that takes a SMPL human motion sequence and produces a Unitree G1 humanoid trajectory" (Section 3.4)
NVIDIA
- Description: Developer of the Cosmos World Foundation Model platform for Physical AI
- Why relevant: Cosmos is cited as a video-based world model approach [70] that WMM positions itself against. NVIDIA's approach operates in pixel space, while WMM argues for SE(3) trajectory space as a more compact and physically grounded alternative. This is a direct competitive positioning against one of the most well-funded world model efforts.
- Quote: "It is also orthogonal to video-based world modeling [3, 9, 70, 92, 103], which captures dynamics implicitly through color changes on a screen, entangling the underlying 3D motion with appearance, lighting, and viewpoint changes" (Section 1)
Physical Intelligence (π0)
- Description: Developer of vision-language-action flow models for robot control
- Why relevant: π0 [8] is cited as a VLA policy approach. WMM's unified approach of predicting SE(3) trajectories directly (rather than actions in a morphology-specific space) represents an alternative architectural philosophy to VLA models that tie action prediction to specific embodiments.
- Quote: "Vision-language-action policies form a complementary set of output interfaces over various tokenizations of action [8, 79, 101]" (Appendix B.1)
4. People Identified
Angjoo Kanazawa
- Lab/Institution: UC Berkeley
- Why notable: Co-author and leading figure in 3D vision and human motion capture. Her prior work on SMPL, 4D reconstruction (Shape of Motion [85]), and articulated templates (GART [47]) directly informs the SE(3) trajectory representation at the core of WMM. She bridges computer vision and robotics, making her lab a key node for physical AI foundation models.
- Quote: Co-authored the paper; the SE(3) representation builds directly on her prior work on 4D reconstruction and motion factorization.
Trevor Darrell
- Lab/Institution: UC Berkeley
- Why notable: Co-author and director of BAIR (Berkeley AI Research), one of the most influential robotics and AI labs. His involvement signals that this work is positioned as a foundational contribution to the physical AI stack, not just a niche benchmark result.
- Quote: Co-authored the paper; BAIR is funded by Meta BAIR partners and the BAIR Humanoid Intelligence Center.
Jiahui Lei
- Lab/Institution: UC Berkeley (primary author)
- Why notable: First author whose prior work on dynamic 4D scene modeling (MoSca [49], MoMaps [48], GART [47], DynMF [43]) directly establishes the theoretical basis for representing complex scene motion as sparse SE(3) trajectory bases. WMM is the logical culmination of this research line — moving from reconstruction to generative prediction.
- Quote: "complex, dense motion of the physical world can be compactly approximated by a sparse set of moving coordinate frames — specifically SE(3) trajectories" (Section 1)
Qianqian Wang
- Lab/Institution: Harvard University
- Why notable: Co-author and developer of Shape of Motion [85], a landmark 4D reconstruction method that produces the SE(3) trajectory labels used to train WMM from raw video. Her work on computing SE(3) trajectories from monocular video is what makes WMM's training data pipeline feasible.
- Quote: Co-authored; her Shape of Motion method is cited as the basis for computing "local frames from these tracks" via "closed-form Procrustes" (Appendix A.4)
5. Operating Insights
A Single Foundation Model Can Replace Multiple Specialized Components in Your Robotics Stack
For a CTO building a robotics platform, WMM's most operationally significant claim is that one trained model serves as policy, world model, inverse dynamics solver, motion planner, and retargeting engine — just by changing the inference mask. From Section 2.3: "A trained WMM is an entire family of conditional samplers under one set of weights." This means you could potentially replace separate perception, planning, and control modules with a single model, reducing integration complexity, training data fragmentation, and deployment overhead. The paper proves this across 6 tasks with SOTA performance, though it notes that joint training across all domains is left to future work (each benchmark currently uses a separate checkpoint).
Robustness to Noisy and Missing Observations Is Built Into the Architecture
Real-world robot deployments suffer from sensor noise, occlusion, and tracker dropouts. WMM handles this natively through two mechanisms: (1) observed tokens can be set to a small positive noise level (e.g., λ=0.05) rather than exactly 0, "which turns each observation into a strong hint and lets the integrator slightly rectify the token" (Section 2.3); and (2) invalid/missing time steps are pinned to pure noise during training and still participate in attention, so the model learns to reason around gaps (Section 2.4). For deployment teams, this means the model degrades gracefully rather than failing catastrophically when perception is imperfect — a critical property for production systems.
Training Data Can Be Sourced from Raw RGB Video at Low Cost
The paper's appendix (Table 12) details the full pipeline for converting raw RGB video to SE(3) trajectory training labels: 131 seconds of processing time and 14.1 GB VRAM for a 40-frame, 32,768-track video at 378×504 resolution. This is remarkably cheap compared to teleoperation or simulation-based data collection. For companies building physical AI datasets, this means internet-scale video could become a viable training source for motion priors — though the current framework still requires depth and estimated states as input, which the authors acknowledge as a limitation.
6. Overlooked Insights
Training on Fewer Trajectories Than Inference Produces Better Results
In the TraceGen experiment (Table 3), the model trained on 128 randomly sampled trajectories (P=128) outperforms the model trained on all 400 trajectories (P=400) on every metric. The authors hypothesize "this stems from the random sampling of 128 out of 400 trajectories during training, which prevents overfitting." This has a profound implication: you may not need to train on the full observation space to generalize to it. For companies with limited compute budgets, subsampling trajectories during training could yield better generalization than full-scale training — a counterintuitive cost-saving strategy.
The Model Cannot Generalize to Novel Morphologies Unseen During Training
Buried in the Limitations section: "The current model does not generalize to novel morphologies unseen during training; further investigation of in-context learning with our proposed context mechanism may suggest solutions in the future." This is a critical caveat for anyone evaluating WMM as a universal robotics foundation model. While the paper demonstrates impressive cross-embodiment results (8+ manipulators, human-to-humanoid), every embodiment was seen during training. A company deploying WMM would still need to collect training data for each new robot platform — the "one model for all robots" vision is not yet realized. The context token mechanism is the authors' proposed path forward, but it remains unproven for zero-shot morphology transfer.