CLAP
“CLAP demonstrates that a single video world model trained across multiple robot embodiments (Franka, WidowX, bimanual YAM, G1 humanoid) can match or surpass models trained exclusively on one robot platform. On the DROID dataset, CLAP-CURR achieves a PSNR of 19.138 and LPIPS of 0.204, compared to the single-embodiment Ctrl-World baseline's 18.928 PSNR and 0.205 LPIPS”
Source→“The robotics community has largely assumed that relative action spaces (e.g., 'move 3cm in x') are easier to learn than absolute action spaces because they have narrower distributions. CLAP's experiments contradict this for end-effector-conditioned video models: 'relative-action spaces underperform absolute-action spaces in future prediction conditioned on end-effector actions across all perceptual metrics and robot environments, e.g., by about 14.6% in LPIPS in the DROID environment'”
Source→“Although state-of-the-art baselines exist in the Bridge environment (e.g., WorldGym, Cosmos-Predict 2.5), we found these baselines to underperform our internal Bridge video models”
Source→“CLAP demonstrates few-shot adaptation to the G1 humanoid's 26-dimensional action space (7-DoF per arm, 6-DoF per hand). The adapted model achieves PSNR 15.151 on G1 video prediction”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.