Ant Lingbo
“Released 6 open-source models in their second generation: Limbo Viren (geometric vision pre-training), Limbo Depth (spatial perception), VLA 2.0 (cross-embodiment action model), Video (MoE-native video generation), Word (causal pre-training), and VA 2.0 (the integrated robot-native foundation model).”
Source→“In 26 this year, or rather this half-year, we made what I think is a fairly big decision: since we can't push the digital-world teams to change, we'll redo everything from scratch based on the physical world's needs.”
Source→“We built our pre-training purely from a geometric angle — you can roughly understand it as points, lines, and surfaces. Starting from edges, we let the model better understand the concept of geometry.”
Source→“The robot data is about two orders of magnitude smaller than internet data. Even with our 60,000 hours this time — honestly compared to last time, it's not an order-of-magnitude improvement.”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.