Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/ARXIV PHYSICAL AI RESEARCH/DepthWorld: 3D World Model for R…
PAPR
// RESEARCH PAPER
ARXIV PHYSICAL AI RESEARCH

DepthWorld: 3D World Model for Robot Manipulation

DATE October 6, 2026SOURCE ARXIV PHYSICAL AI RESEARCHPARTICIPANTS JAI BARDHAN, JOSEF SIVIC, VLADIMIR PETRIKARXIV 2610.08780
// SUMMARY

Investor & Operator Brief — Physical AI

Bottom line: This paper does two things that matter commercially: (1) it turns the largest open teleoperation dataset (DROID, ~71k episodes) into a properly calibrated 3D dataset at millimeter accuracy, and (2) it shows that adding depth supervision to a video world model improves the RGB video itself by +1.48 dB PSNR — without touching the pretrained backbone. If you're building or evaluating world-model companies, both findings change how you should think about data assets and architecture.


1. Key Themes

Retroactive 3D Calibration of Existing Robot Data at Scale

The core data contribution is a pipeline that takes any multi-view stereo teleoperation dataset with a known URDF and produces supervision-quality 3D annotations — dense metric depth plus camera extrinsics anchored to the robot's physical base frame. Applied to DROID, this yields "DROID-3D—a calibrated 3D corpus providing dense metric depth and recalibrated multi-view extrinsics for over 70,000 episodes" (Section 1, Contribution 1). The key trick is a joint factor graph that pools all episodes from the same physical robot to jointly solve for per-scene camera corrections and shared robot kinematic parameters (hand-eye mount, joint encoder offsets). The results are dramatic: "23× improvement on EE, 13× on WE, 8× on RD, and +37 absolute points on IoU" over a PointWorld-style per-scene baseline (Section 5, Table 1), and against real ground truth measured with a ChArUco board in their own lab, "our approach recovers extrinsics to 4.9 mm/0.39° median error and hand-eye to 3 mm/0.3°... joint offsets are recovered to 0.14° mean error across all seven joints" — versus the baseline's "72 mm / 3.0° median error, an order of magnitude worse" (Section 5).

Depth Supervision Improves RGB Prediction — For Free

The headline empirical finding: training a video world model to jointly predict depth alongside RGB makes the video better, not just the geometry. "At equal training budget, DepthWorld gains +1.48 dB PSNR over an RGB-only baseline of identical architecture and data" (Abstract; confirmed in Table 2: external PSNR 22.63 → 24.09, wrist 16.98 → 17.96). This is a counterintuitive result — adding a second prediction target with the same data and compute improves the first target. Practically, it means world-model companies currently training RGB-only models are leaving quality on the table.

Spatial Latent Tiling: Adding Modalities Without Destroying Pretrained Priors

The architectural contribution is a way to bolt depth onto Stable Video Diffusion without re-initializing any pretrained weights. Instead of expanding VAE input channels or running a parallel depth U-Net — both of which "require destructive weight re-initializations" that "risk corrupting the strong visual and motion priors" (Section 1) — they exploit the fact (from Marigold) that the VAE already encodes depth maps faithfully, and simply tile RGB and depth side-by-side in a wider latent grid (72×80, Figure 3). "The only network change is extending the spatial position embeddings to cover the wider grid" (Section 4.1). The ablation in Appendix C.2 (Table A2) shows tiling beats both dual-branch and channel-expansion architectures on nearly every metric (+1.7 dB external PSNR over dual-branch).

Geometrically Coherent Rollouts Enable Real Policy Evaluation

The paper's framing is that RGB-only world models "produce visually plausible rollouts whose geometry is internally incoherent: predicted depth disagrees across views of the same scene, the model struggles with geometric reasoning about occlusion and contact" (Section 1). Figure 4 shows a concrete failure mode with real deployment consequences: "Without depth supervision, the RGB-only world model cannot distinguish the open wardrobe door from a solid surface and simulates the arm colliding with it. DepthWorld accurately generates the arm going inside the wardrobe." If you're using world models to evaluate or improve policies, a model that hallucinates collisions where none exist will give you garbage signal.

Competitive Positioning Against Other 3D World Models

DepthWorld outperforms the two comparable systems. Against an adapted TesserAct at equal training budget: "external PSNR 19.43 vs 24.07 (ours), AbsRel 0.154 vs 0.074 (ours)" (Section 5). Against PointWorld (NVIDIA's point-cloud world model), the comparison is nuanced: PointWorld is more accurate on moving objects at short horizon (23 vs 31 mm at 2s), but DepthWorld degrades far less over long rollouts — "whole scene PW 8→18 mm vs. ours 8→9 mm; moving region PW 21→58 mm vs. ours 25→32 mm" (Section 5, Table A5). DepthWorld's rollouts stay coherent over 8 seconds; PointWorld's point clouds drift and scatter.


2. Contrarian Perspectives

"RGB-Only Video Priors Are Geometrically Broken — and Everyone Building World Models Knows It"

The paper directly challenges the current wave of RGB-first world-model companies (and the video-generation-as-world-model thesis generally). The claim is that frame-by-frame photorealism is a misleading metric: rollouts "look correct frame-by-frame but do not compose into a consistent 3D world" (Abstract). The evidence is concrete — Figure 1 shows the RGB baseline losing the shape of cutlery entirely mid-rollout, and Figure 4 shows it simulating a collision with an open door. The contrarian implication: if your world model is RGB-only, its rollouts cannot be trusted for the three use cases everyone cites (policy evaluation, improvement, planning), because "the geometric fidelity of the rollout determines whether the signal is usable" (Section 1).

"Don't Fine-Tune the Backbone Architecture — Most 'Multi-Modal' Modifications Make Things Worse"

The conventional engineering instinct for adding a modality is to expand input channels or add a parallel branch. This paper argues both approaches are actively harmful to pretrained video models, and shows data: spatial tiling beats dual-branch and channel-expansion on "every RGB metric on both views" (Appendix C.2, Table A2). The reasoning: "Conventional architectural modifications for incorporating spatial modalities... require destructive weight re-initializations. These interventions risk corrupting the strong visual and motion priors that make these foundation backbones effective in the first place" (Section 1). For any team fine-tuning SVD-class backbones, this is a cheap, validated architectural pattern.

"Per-Scene Calibration Is the Wrong Abstraction — Calibrate the Robot, Not the Scene"

Most robotics data teams treat each recording session as an independent calibration problem. This paper argues that's fundamentally flawed: "Optimizing scenes independently simply absorbs these hidden, systematic errors into the per-scene extrinsics" (Section 3.2). By pooling all episodes of a robot (~1,000 to ~9,000 per robot, per Appendix B.2) and jointly solving for shared kinematic parameters, they "multiply the correction signal for the latter by the number of episodes" (Section 3.2). The proof: the per-scene baseline "fails to converge on ∼18% of scenes (180 of 995); our coupled optimization has no analogous per-scene failure mode" (Section 5). The broader contrarian point for data-heavy robotics companies: your factory calibration and joint encoders have systematic biases you can recover for free from data you already have.


3. Companies Identified

Stability AI — Provider of Stable Video Diffusion (SVD), the pretrained backbone underlying DepthWorld and Ctrl-World. Why relevant: SVD is becoming the de facto open video prior for robot world models. Quote: "DepthWorld, a Stable Video Diffusion-based world model that jointly predicts multi-view RGB and depth via spatial latent tiling, leaving the pretrained Variational Autoencoder (VAE) unchanged" (Abstract).

Stereolabs — Maker of the ZED 2 (external) and ZED Mini (wrist) stereo cameras used across all of DROID. Why relevant: their SDK depth is shown to be inadequate for this domain — "Classical semi-global matching (SGM) algorithms yield metric scale but fail on the texture-poor tabletops and specular surfaces of the robot arm" (Appendix A.1), and training on ZED ULTRA depth actively degraded the world model. Their hardware is ubiquitous in teleoperation rigs, but their stock depth output is not supervision-quality.

Franka (Panda arms) — The robot embodiment for all of DROID. Quote: "The recordings come from 13 institutions on Franka Panda arms, each with two table-mounted ZED 2 external cameras and a wrist-mounted ZED Mini" (Section 3.3).

DROID (Stanford-led consortium) — The dataset itself, 71,100 usable episodes across 28 robots and 13 labs. Why relevant: DROID-3D transforms this from an RGB dataset into a 3D dataset, materially raising its value to anyone training geometric models. Quote: "DROID is unique among real-world manipulation datasets at this scale in providing synchronized multi-view stereo with known baselines and a publicly released URDF" (Section 2).

NVIDIA — PointWorld, the main 3D world-model competitor, is a NVIDIA-led effort (Fei-Fei, Fox, Mousavian, Liu per reference [34]). Why relevant: it's "the only action-conditioned 3D world model trained on DROID" (Section 5) and serves as the calibration baseline. The comparison shows PointWorld's point-cloud representation drifts over long rollouts while DepthWorld stays flat — but the appendix honestly notes "Against raw depth, PointWorld's point representation retains a large advantage on moving objects" (Appendix D.4).

Google DeepMind / Google — RT-1 and Open X-Embodiment are cited as datasets that "carry geometry as a byproduct of the capture pipeline: monocular RGB in Open X, single RGB-D in Bridge" (Section 2) — i.e., competitors' data assets that lack the stereo + URDF structure that made DROID-3D possible.

RoboCasa (NVIDIA) / simulation dataset providers — Cited as carrying "perfect geometry but inherit the sim-to-real gap that motivates training world models on real data in the first place" (Section 2).

Ctrl-World (Stanford/Chelsea Finn lab) — The SVD-based world model that forms DepthWorld's direct backbone: "Ctrl-World [1], an SVD-based controllable world model trained on DROID [18] forms our backbone" (Section 2). DepthWorld is essentially Ctrl-World + depth, retrained on the same data — making the +1.48 dB comparison a clean apples-to-apples result.

TesserAct (MIT-IBM lineage, per reference [4]) — RGB-D-normal video world model, adapted and beaten by ~4.6 dB PSNR and ~2× AbsRel at equal budget (Section 5, Appendix D.5).


4. People Identified

Jai Bardhan — Czech Institute of Informatics, Robotics and Cybernetics (CIIRC), Czech Technical University in Prague; lead author. Notable: also an author of PersistWorld (reference [2], RL-stabilized world models) and REALM (reference [42]) — this group is building a coherent world-model research program, and this is their second major world-model paper. Quote (from Limitations): "recent works like PersistWorld [2] have shown that RL post-training can significantly improve rollout stability."

Josef Sivic — CIIRC CTU Prague; senior author, one of the most influential computer vision researchers of the past two decades (co-inventor of the visual bag-of-words, work foundational to visual SLAM and place recognition). His involvement signals that the calibration pipeline is serious classical-vision engineering, not just deep learning.

Vladimir Petrik — CIIRC CTU Prague; senior author, robotics-focused, co-author of REALM and AlignPose (reference [51]). The group is EU-funded (Horizon Europe AGIMUS, euROBIN, ERC FRONTIER — Acknowledgments), meaning this work is open and reproducible rather than locked inside a corporate lab.

Chelsea Finn (Stanford) — Author of Ctrl-World (reference [1]), the backbone world model. The Stanford–CTU lineage matters: the base architecture the community is converging on for manipulation world models comes from this line of work.

Fei-Fei Li, Dieter Fox (NVIDIA) — PointWorld authors (reference [34]). NVIDIA is investing in 3D world models for manipulation; this paper is the first credible open head-to-head against their approach.

Abby Khazatsky, Karl Pertsch, et al. (DROID consortium) — Authors of the DROID dataset (reference [18]); the data asset this entire paper is built on.


5. Operating Insights

Your Existing Teleoperation Data Is Worth More Than You Think — If It Has Stereo + URDF

The calibration pipeline runs on data you may already have sitting in cold storage. It recovers hand-eye calibration to 3mm, joint encoder offsets to 0.14°, and camera extrinsics to 4.9mm — all post-hoc, from teleoperation episodes alone, at "roughly 12.5 minutes per robot group" of solver time (Section 3.2). If you're a company with a fleet of robots collecting stereo teleoperation data, this is a playbook for converting that data into a 3D supervision asset without new collection. Conversely, if you're diligencing a robotics data company, ask whether their rig has known stereo baselines and a public URDF — that's what made DROID liftable and what makes monocular-RGB datasets (Open X-Embodiment) much harder to upgrade.

Add Depth as a Joint Prediction Target via Latent Tiling — It's a Free Win on RGB Quality

For any team fine-tuning SVD-class video models for robotics: the spatial latent tiling pattern (encode depth as a grayscale image, tile it next to RGB in the latent grid, extend position embeddings only) delivers +1.48 dB PSNR on RGB at identical data and compute, with zero re-initialization of pretrained weights. The two-stage training schedule (40k steps RGB+depth only, then 50k steps with the point-map head at λpm = 0.005, Appendix C.6) and the depth preprocessing recipe (log transform, 95th-percentile clip, z-buffer min-pooling — Appendix C.1) are directly copyable engineering details. Total cost: "approximately two days on two H200 nodes" (Appendix C.6) — this is not a frontier-compute result.

Audit Your Auxiliary Signal Quality Before Adding It

The depth-source ablation (Appendix C.3, Table A3) contains a warning every ML lead should internalize: training on sparse/noisy ZED ULTRA depth degraded RGB prediction by 2.0 dB on external views "even though the RGB supervision is unchanged." Quote: "Sparse and noisy depth is therefore not a neutral auxiliary signal but actively harms the joint RGB–depth model." Bad auxiliary supervision is not harmless regularization — it's poison. Vendors selling "depth annotations" for robot data should be evaluated on boundary sharpness and coverage in exactly the textureless/specular conditions that dominate manipulation scenes.


6. Overlooked Insights

The Latent-Space Depth Representation Has a Hard Error Floor of ~45mm on Moving Objects

Buried in Appendix D.4: the SVD VAE's encode-decode round trip alone "displaces the ground truth by about 45 mm (median)" on moving points — an error floor "its network cannot go below." The authors are candid that "Against raw depth, PointWorld's point representation retains a large advantage on moving objects, a representational cost of latent-space depth prediction that we consider an important direction for future work." For anyone planning to read precise 3D geometry out of a latent-space world model (e.g., for contact-rich manipulation or grasp planning), sub-5cm accuracy on moving objects is currently a representational ceiling, not a training problem. This materially constrains what "world model as simulator" claims can deliver today for fine manipulation.

The Evaluation Measures Same-Robot Generalization — Not the Thing Customers Actually Need

A quiet admission in Appendix D.3: "the held-out trajectories come from the same robots and labs seen in training; these numbers therefore measure held-out-trajectory prediction quality, not cross-embodiment or cross-lab generalization." The +1.48 dB result and the depth metrics are within-distribution. Anyone extrapolating these numbers to a world model that generalizes to a new customer's kitchen, new robot, or new camera placement is over-reading the evidence. Relatedly, wrist-view depth remains substantially worse than external-view depth (δ1 ≈ 0.82 vs > 0.94, Section 5) due to "rapid motion, proximity to the manipulated object, and frequent self-occlusion" — and the wrist camera is precisely the view closest to the manipulation action, which is where geometric fidelity matters most for grasping.