Stanford University research group behind the OpenVLA open-source vision-language-action model.
“On the CoffeeServeMug task, this pipeline achieves 90% success rate, outperforming Diffusion Policy, ACT, and SmolVLA baselines — despite the video model never seeing this specific task during training”
Source→“The model trained only on forward examples (robot motion → scene) generalizes zero-shot to the inverse direction (object motion → robot), which the authors note was unexpected”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.