Mengya Liu
Lead author of the LARA framework for Vision-Language-Action models.
“LARA: Latent Action Representation Alignment for Vision-Language-Action Models”
Source→“LARA's core bet is that you can squeeze more performance out of existing unlabeled video (human or robot) by jointly training a Latent Action Model (LAM) and a diffusion-based VLA policy, rather than treating them as separate sequential steps.”
Source→“For GR00T-N1.6, this yields improvements of +1.3% on SIMPLER-ENV and +5.56% on real-world G1 humanoid tasks (Table 1, Table 2).”
Source→“LARA is applied to π0.5 as a post-training module, achieving +0.9% improvement on LIBERO average.”
Source→“LARA achieves ~30% average improvement when adapting from OXE-pretrained models to entirely new embodiments (Unitree G1 humanoid, GR1-Sim) not seen during pretraining.”
Source→“LARA (full) outperforms the best LAM pseudo-label baseline (Moto-GPT) by +16.8% on SIMPLER-ENV (Table 1) while using only OXE-constrained data.”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.