“N0-TWAM is the first world-action model that predicts future touch jointly with future vision, rather than treating tactile data as a side input. It was pre-trained on "tens of thousands hours of real-robot data… spanning six embodiments and 450 tasks"”
Source→“On the UniVTAC benchmark, "the vision-only world-action baselines trail even the VLA policies here (LingBot-VA 31.4, FastWAM 48.0, GigaWorld-Policy 16.5): predicting the future scene is not enough for tasks defined by contact"”
Source→“π0.5 is the strongest baseline. As a large vision-language-action policy it brings broad semantic priors and dependable visually guided placement… but it regresses actions directly from the current frame, with no model of the next moment and no touch, so it can neither anticipate nor feel a contact event”
Source→“The pipeline uses "Gemini 3.5 Flash to generate a short sub-task instruction for each clip, with human spot checks"”
Source→“Training runs on 128 NVIDIA H800 GPUs under FSDP2 with bf16 mixed precision”
Source→“The Wan2.2-TI2V-5B video-diffusion transformer family provides the pretrained video expert that N0-TWAM's backbone is warm-started from”
Source→“Listed under "Project Lead" and "Pre-Training" in the Contributors section”
Source→“Listed under "Project Lead", representing the academic side of the collaboration”
Source→“The most ubiquitous contributor — listed in Pre-Training, Post-Training (Simulation), Post-Training (Real Robot), and Data Processing”
Source→“Listed under "Project Lead", likely representing the industry/commercial side of the collaboration”
Source→“Listed under "Project Lead", likely representing the industry/commercial side of the collaboration”
Source→“N0-VTLA is, by its own claim, the first vision-tactile-language-action model pretrained on tactile data at scale. The result: on a 20-task simulation suite, N0-VTLA reaches 63.8% mean success against 44.0% for the strongest baseline (π0.5), and wins all nine real-robot NeoReal tasks”
Source→“The pretraining corpus, NeoData, spans multiple robot platforms (ARX X5, UR5e, Flexiv, Franka, Piper) and a handheld UMI-style gripper, with every gripper finger carrying a self-developed visuo-tactile sensor”
Source→“The core architectural insight is that touch should condition actions as a *prediction* of future contact, not as a current observation... a lightweight predictor emits latent tactile tokens z that 'estimate the net tactile change over the coming action chunk'... After Stage 1 training, these latents retrieve their matching future-tactile target at 92.3% top-1 accuracy vs. 3.2% chance”
Source→“NeoteAI / Fudan TEAI: The authors' organizations, developers of N0-VTLA, NeoData, NeoReal benchmark, and the self-developed visuo-tactile sensor. They have built the full stack: sensor hardware, data pipeline, model architecture, training recipe, and benchmarks”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.