Zuxuan Wu
Prominent researcher in efficient deep learning and video understanding at Fudan University.
“Listed under "Project Lead", representing the academic side of the collaboration”
Source→“N0-VTLA is, by its own claim, the first vision-tactile-language-action model pretrained on tactile data at scale. The result: on a 20-task simulation suite, N0-VTLA reaches 63.8% mean success against 44.0% for the strongest baseline (π0.5), and wins all nine real-robot NeoReal tasks”
Source→“Zuxuan Wu: Project Lead and likely faculty advisor, given Fudan affiliation”
Source→“Tianyi Lu, Hui Zhang, Zijie Diao, Junke Wang, Shengqi Xu, Xing Lin, Guojin Zhong, Ziyi Ye, Peng Wang, Zuxuan Wu, et al.”
Source→“VLA-Pro stores task-specific LoRA adapters as parameterized procedural memories during training. At inference time, VLA-Pro retrieves relevant procedural memories based on the current multi-modal context and dynamically fuses these memories for generating the current action chunk.”
Source→“His involvement suggests the efficiency angle of LoRA-based memory (small, modular adapters vs. full fine-tunes) is a deliberate architectural choice informed by efficient ML principles, not just a pragmatic shortcut.”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.