Hongyu Ding
Hongyu Ding is a researcher affiliated with Xi'an Jiaotong University and Peking University, where he collaborates with advisors including Jian Cheng, Yang Gao, and Jiebo Luo. He is best known as the lead author of Uni-LaViRA, a unified agentic architecture for embodied navigation that leverages pretrained multimodal large language models to perform vision-language navigation, object navigation, embodied question answering, and UAV navigation across heterogeneous robot platforms in a zero-shot manner. His research focuses on embodied AI, navigation foundation models, and deploying multimodal language models on physical robots.
“Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation”
Source→“for navigation (as opposed to contact-rich manipulation), the action space already falls within the natural output manifold of pretrained MLLMs. Directional commands ("turn left") and pixel-level visual targets (bounding boxes) are things foundation models already produce reliably from pretraining. This means navigation can be *reasoned* by an agent rather than *learned* from robot trajectories.”
Source→“Qwen3.5-27B runs for VA in simulation and on the wheeled platform; Qwen3.5-9B-Q4 runs locally on the Unitree G1 and Go1's Jetson Orin NX.”
Source→“Agilex (Cobot Magic) — Wheeled bimanual platform used for real-world VLN-CE and ObjectNav validation. Runs Qwen3.5-27B-Q4 locally on an RTX 4090 with no remote API calls.”
Source→“The Go1 is noted as "the most agile of the four platforms and executes wide turns without re-planning." The G1 experienced calibration drift of ~2cm after 30 minutes, causing overly wide turns at doorways.”
Source→“Habitat-Sim is used for all ground robot benchmarks. SAM and Grounding DINO are mentioned as future tools for the Vision Action Model when bounding-box confidence is low.”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.