Kechen Liu
“CLAP demonstrates that a single video world model trained across multiple robot embodiments (Franka, WidowX, bimanual YAM, G1 humanoid) can match or surpass models trained exclusively on one robot platform. On the DROID dataset, CLAP-CURR achieves a PSNR of 19.138 and LPIPS of 0.204, compared to the single-embodiment Ctrl-World baseline's 18.928 PSNR and 0.205 LPIPS”
Source→“The robotics community has largely assumed that relative action spaces (e.g., 'move 3cm in x') are easier to learn than absolute action spaces because they have narrower distributions. CLAP's experiments contradict this for end-effector-conditioned video models: 'relative-action spaces underperform absolute-action spaces in future prediction conditioned on end-effector actions across all perceptual metrics and robot environments, e.g., by about 14.6% in LPIPS in the DROID environment'”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.