arXiv Physical AI
“This produces 277 hours of dynamically feasible, contact-consistent trajectories across six locomotion task families... with 15,276 total episodes.”
Source→“WOLF-VLA: Whole-Body Humanoid Optimal Locomotion Framework for Vision-Language-Action Learning”
Source→“The core idea is to distill expert behavior, obtained from reinforcement learning, teleoperation, or other controllers, into training data that can be used to fine-tune compact VLA models... reducing manual system integration and lowering the barrier for deploying new robot behaviors.”
Source→“Patch Policy: Efficient Embodied Control via Dense Visual Representations”
Source→“By directly consuming dense patch tokens from Vision Transformers (ViTs) rather than compressing them into a single global vector, Patch Policy achieves a "40% relative improvement over policies using state-of-the-art global-pooled representations"”
Source→“Patch Policy "surpasses fine-tuned OpenVLA-OFT by 18% while using roughly 0.7% of the parameters"”
Source→“ICWM treats it as a system identification problem solved through in-context learning”
Source→“ICWM enables robot policies to autonomously infer essential system variables from a short history of self-generated, task-agnostic interactions.”
Source→“Participants: Tianyi Lu, Yu-Gang Jiang, et al. (arXiv Physical AI)”
Source→“ENPIRE framework provides the necessary abstraction: 'reset the scene, execute a policy, verify the outcome, and refine the next iteration' (Abstract). This transforms real-world robot learning into a 'controllable optimization procedure' (Abstract)”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.