RoboJEPA: Scaling Robotic Latent World Models
- 01Predictable Scaling Laws for Robotic World Models
- 02Capabilities Emerge in a Predictable Order at Specific Compute Thresholds
- 03Offline "Imagination Error" Is a Reliable Proxy for Expensive Real-Robot Evaluation
- 04Zero-Shot Real-Robot Control from a Single Goal Image
- 05Training Must Match Deployment Horizon
Bottom line: Meta FAIR just did for robotic world models what Kaplan/Chinchilla did for LLMs: established predictive scaling laws. This means robotics capability planning can now be forecast before spending compute — and it's the strongest public evidence yet that world-model planning is a credible alternative (and complement) to VLA policies.
1. Key Themes
Predictable Scaling Laws for Robotic World Models — Robotics Becomes a Forecastable Engineering Discipline
The core contribution is a fitted scaling law — L(C) = E + A·C^(α−γ ln C), a "second-order power law" — that predicts world-model prediction error as a function of training compute, and critically, extrapolates accurately to models 4× larger than those used to fit it. From Section 3.1: "The second-order power law extrapolates best, with a held-out error of 0.6 and 1.4×10⁻³ on DROID and RoboCasa respectively, roughly 2–3× lower than the standard power law." The authors frame the strategic value directly in the Introduction: "knowing which capabilities a world model will have at a given training-compute budget lets practitioners decide whether to invest in a larger model, more demonstrations, or more compute, and reveals when a data budget has been saturated." This is the first time a robotics team can answer "what do I get for another 10× compute?" with a number instead of a hope.
Capabilities Emerge in a Predictable Order at Specific Compute Thresholds
The paper maps a capability ladder tied to compute budgets (Section 3.2–3.3): "first the 3D control of the end-effector [~10²⁰ FLOPs], then manipulating a grasped object [~3×10²⁰], then the geometry of the static scene [~10²¹], and finally the dynamics of objects in the scene [~10²²]." Notably, "no model achieves a positive success rate [on object pushing] until around 10²² FLOPs, after which all models begin to succeed, suggesting that fine-grained understanding of surrounding objects emerges after certain training computation regardless of the model size." For anyone building a manipulation product, this is a capability roadmap: object-interaction reasoning is a ~10²² FLOP problem, full stop.
Offline "Imagination Error" Is a Reliable Proxy for Expensive Real-Robot Evaluation
The paper shows the world model's latent rollout error (measurable offline, cheaply) strongly correlates with real-robot planning success. From the Conclusion: this "makes the offline imagination error a reliable proxy for real-robot evaluation and spares much of its computational cost." This matters operationally because real-robot eval is the single biggest cost sink in robotics development — the paper ran "over 50,000 evaluation episodes" across two platforms to establish this. A company could now gate hardware trials on an offline metric.
Zero-Shot Real-Robot Control from a Single Goal Image — No Task-Specific Training
RoboJEPA deploys directly on a real Franka via planning toward one goal image, with no task-specific fine-tuning, using "the same set of hyperparameters across embodiments and tasks" (Section 2.2). At 8B parameters it achieves 67% success on Grasp, 50% on Object Lift, and 27% on Pick and Place (Table 3). The 8B model is "the largest JEPA predictor model trained to date," trained on 23 datasets spanning 12 embodiments and 15,022 hours of video (6,692 hours action-synchronized) — a demonstration that heterogeneous multi-embodiment data can be unified into one model.
Training Must Match Deployment Horizon — Mismatch Breaks Scaling Predictability
A buried but important finding (Section 3.3): models trained with short 2-step rollout prediction showed no clean scaling frontier on long-horizon tasks — "smaller models often beat larger ones (for example, the 300M planner reaches a higher final reward than the 1B one)." Only after retraining with 10-step rollouts during the cooldown phase did predictable scaling emerge for long-horizon planning. The fix was cheap: "the long flat-learning-rate stage only needs to learn the coarse dynamics cheaply, and the computationally expensive long-rollout, high-resolution signal is spent where it matters, in the final annealing phase."
2. Contrarian Perspectives
World Models Beat VLA Policies on Tasks Where Language Data Is Thin
The most commercially interesting result challenges the VLA-first orthodoxy dominating robotics investment. On the real Franka, Physical Intelligence's π0.5 scored 5% on Grasp and 0% on Object Lift, while RoboJEPA-8B scored 67% and 50% respectively (Table 3). The authors' explanation: "verbs such as 'lift' are heavily underrepresented in the DROID dataset... Due to the planning nature of the agent and the task specification as an image goal, RoboJEPA models are not susceptible to this type of degradation." The strategic implication: VLA policies inherit the long-tail coverage problems of their text-action training data, while goal-image planning sidesteps language entirely. (Caveat the authors acknowledge: these are "contextual references rather than directly comparable baselines" due to different goal specifications and training.)
The Bottleneck Is Data, Not Parameters — and the Field Is Already Saturated
Against the "just scale it" narrative, the paper's own scaling fits show RoboJEPA is near data saturation: "the two best candidates... estimate nearly the same irreducible error, ≈0.2 on DROID and ≈0.17 on RoboCasa, indicating that RoboJEPA is close to saturation on both evaluations under its current training and data budget" (Section 3.1). The Conclusion is blunt: "Because we train for multiple epochs over a fixed corpus, our fits already sit close to data saturation, so pushing the frontier further will likely require more and more diverse interaction data rather than parameters alone." For investors, this says the scarce asset in physical AI is proprietary interaction data — the 15,000-hour public corpus is already tapped out.
Latent Prediction Beats Pixel Generation for Scaling World Models
The paper implicitly argues against the video-generation-as-world-model camp (Genie, GAIA-style models): "RoboJEPA predicts in the representation space of a frozen V-JEPA encoder, which is what allows us to scale the predictor to billions of parameters on robot data without paying the computational cost of pixel-level reconstruction at every step" (Section 4). Pixel-space generative world models are positioned as simulation/visualization tools, while latent models are the scalable path to control. The diffusion decoder in this paper is used "purely as a visualization tool" — pixels are for humans, latents are for planning.
3. Companies Identified
Meta (FAIR) — Primary author affiliation; built RoboJEPA, V-JEPA 2/2.1, and released all checkpoints and deployment code ("We release all model checkpoints together with our training and robot-deployment code"). Signals Meta is investing seriously in world-model-based physical AI as an alternative to the VLA stack.
Physical Intelligence — Makers of π0, π0-FAST, and π0.5, used as the VLA reference models on the real Franka. Their models "work best on Pick and Place, the hardest task, but degrade significantly on the others" (Section 3.4) — a direct competitive data point showing VLA fragility outside well-represented task language.
NVIDIA — The diffusion decoder is built on the Cosmos tokenizer ("Cosmos-CV4x8x8", Section 2.3); NVIDIA's Cosmos is cited as releasing "physical-AI world foundation models at industrial scale" (Section 4). NVIDIA's infrastructure layer is embedded in competitors' stacks.
1X Technologies — Their humanoid world-model dataset (23.5k episodes, 90 hours) is part of the training mixture (Table 1). Their data is now fueling a Meta world model.
AgiBot — AgiBot World is the single largest data source in the mixture: "167.5k episodes, 8,036 video hours, 2,734 action hours" of dual-arm and humanoid data (Table 1). AgiBot's data-collection strategy is quietly becoming foundational infrastructure.
Wayve — Cited for the GAIA series of driving world models "scaled to billions of parameters," though "technical details beyond GAIA-2 are not publicly available" (Section 4). A closed-source competitor in world-model scaling for autonomy.
Dyna Robotics — Cited for Dyna-2, which "scales a world-action model over a million hours of human and robot video" (Section 4). The paper distinguishes itself: Dyna-2 measures action-prediction quality, while RoboJEPA fits "an explicit scaling law" and connects it to downstream planning.
Google DeepMind — Genie 2 cited as a "large-scale foundation world model" for interactive environments, positioned outside robot control (Section 1).
Hugging Face — LeRobot community dataset (SO-101 embodiment, 21.5k episodes) is in the training mixture (Table 1).
Franka — The hardware platform for all real-robot deployment ("a single-arm Franka robot and two cameras," Section 3.4).
4. People Identified
Yann LeCun — Meta/NYU; architect of the JEPA paradigm this work instantiates ("JEPAs were proposed as a path towards self-supervised models that learn abstract representations of the world by predicting in a latent space," Section 4). His world-model thesis now has scaling-law evidence in robotics.
Jeannette Bohg — Meta FAIR / Stanford; co-senior author and a driving force behind the DROID dataset used for both training and real-robot deployment. Central figure connecting large-scale robot data collection to foundation-model research.
Mahmoud Assran — Meta FAIR; joint last author and lead of V-JEPA 2, the action-conditioned predecessor RoboJEPA builds on. The through-line from V-JEPA 2-AC's "zero-shot Franka manipulation by latent planning" (Section 4) to this paper's scaled version is his research program.
Artem Zholus — First author (Meta/Mila/Chandar Research Lab/Polytechnique Montréal); also co-authored work on "what makes a latent space useful for robotic world models" (Nilaksh et al., cited in references). A name to watch in latent world models.
Sarath Chandar — Mila - Quebec AI Institute / Chandar Research Lab; co-author representing the Canadian academic axis of the collaboration.
Chelsea Finn & Sergey Levine (referenced) — Stanford and Berkeley respectively; authors of the π0/π0.5 VLA baselines and much of the underlying data ecosystem (DROID, Open X-Embodiment, Bridge). Their behavior-cloning school is the incumbent approach this paper benchmarks against.
Nicolas Ballas & Adriana Romero Soriano — Meta FAIR; joint last authors, core V-JEPA team members.
5. Operating Insights
Build Offline Evaluation Gating Before Building Hardware Programs
The demonstrated correlation between offline imagination error and real-robot success means a CTO can now run cheap, large-scale offline evals (the paper's L1 rollout metric on held-out data) and only spend robot time on models that clear a threshold. The authors' own recommendation (Section 3.2): "for building robotic world models, we recommend first building downstream evaluations that reflect the desired capabilities, then estimating the minimal training compute required to achieve them." Budget hardware eval as validation, not exploration.
Align Training Rollout Length With Your Deployment Horizon — Cheaply
The K=2 vs. K=10 lesson (Section 3.3) is a direct playbook item: pretrain with short rollouts at low cost, then spend expensive long-rollout, high-resolution training only in the final annealing phase. This "long-horizon cooldown" turned a noisy, non-monotonic compute-to-performance relationship into a clean scaling frontier — and it's a recipe any team can apply without retraining from scratch.
Planning-Based Control Has an Inference Infrastructure Bill You Must Budget For
CEM planning requires "thousands of forward passes of the world model" per decision (Section 2.2). The team needed a rolling KV-cache, a custom precompiled inference engine (Appendix I), and a direct GPU-cluster connection to the robot — and even then, "while it is not real-time, it takes only several seconds to perform a multi-step planning even for an 8B model." Anyone productizing world-model planning needs to plan for GPU serving infrastructure at the edge or near-edge, and for control frequencies measured in seconds, not milliseconds. This is viable for manipulation; it is not yet viable for dynamic tasks.
6. Overlooked Insights
The Encoder Is a Hidden Single Point of Failure — V-JEPA 2 Didn't Scale, V-JEPA 2.1 Did
Buried in Section 2.1: "We use V-JEPA 2.1 as the encoder because early in the work we found that it leads to successful and predictable scaling of the world model, whereas V-JEPA 2 does not (we hypothesize why in Appendix K)." The entire scaling-law result is conditional on the representation underneath the dynamics model. For any team building on frozen foundation encoders, this is a warning: your world model's scalability is hostage to encoder quality, and the paper's own limitation section notes the laws "describe a dynamics model on top of a fixed representation rather than the encoder and predictor jointly."
Tiny Models Already Work on Real Robots — the Floor Is Lower Than Expected
Figure 10's caption notes that on real-robot deployment, "all RoboJEPA models, including very small models, demonstrate non-zero success rates" — a 22M-parameter model grasps real objects at non-trivial rates. Combined with the finding that basic end-effector control emerges at ~10²⁰ FLOPs (trainable on a modest budget), this suggests entry-level world-model planning for simple tasks is accessible to startups without frontier-scale compute — the expensive thresholds only kick in for object interaction and long-horizon reasoning. Also worth noting from the architecture appendix: "We found QK-normalization to be of particular importance for training stability when scaling the dataset size and the model size" (Section 2.1) — a small detail that reportedly determines whether large-scale training runs survive at all.