Yiran Qin
Yiran Qin is a Ph.D. student at The Chinese University of Hong Kong, Shenzhen, and a visiting Ph.D. student at Oxford University. Her research focuses on embodied AI, multimodal large language models, and world models, with publications including MP5, a multi-modal open-ended embodied system in Minecraft, and a survey of interactive generative video. She has also published work on supervised LiDAR-camera fusion for 3D object detection and video generation models as world simulators.
“Listed under "Project Lead" and "Pre-Training" in the Contributors section”
Source→“N0-VTLA is, by its own claim, the first vision-tactile-language-action model pretrained on tactile data at scale. The result: on a 20-task simulation suite, N0-VTLA reaches 63.8% mean success against 44.0% for the strongest baseline (π0.5), and wins all nine real-robot NeoReal tasks”
Source→“Yiran Qin: Project Lead for N0-VTLA. Listed as contributor across pretraining, offline RL, and project leadership”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.