“training-time RTC simulates this latency during post-training and teaches the model to generate future actions conditioned on already committed actions (Section 3.3, citing Black et al. [5])”
Source→“DeMaVLA adopts Qwen3-VL [1] as the VLM backbone (Section 3.2)”
Source→“Taiyi Su... Jian Zhu... Yi Xu... AIRC, Midea Group, Tongji University (Author affiliations)”
Source→“Taiyi Su... Jian Zhu... Yi Xu... AIRC, Midea Group, Tongji University (Author affiliations)”
Source→“Taiyi Su... Jian Zhu... Yi Xu... AIRC, Midea Group, Tongji University (Author affiliations)”
Source→“DeMaVLA introduces a single Vision-Language-Action (VLA) model capable of folding multiple types of garments (shirts, skirts, pants, towels) from random initial states”
Source→“DeMaVLA moves beyond category-specific folding policies by using a single checkpoint to handle multiple household folding tasks with different garments, initial states, and long-horizon bimanual routines (Section 5)”
Source→“We compare DeMaVLA with a state-of-the-art VLA baseline... π0... We fine-tune it on our folding tasks with the same training-time RTC setting (Section 4.2)”
Source→“On the real-world folding benchmark, DSWAM achieves 96.3% average success rate versus DeMaVLA's 92.5%, while reducing completion time from 2'18" to 1'44"”
Source→“BF16 TensorRT reduces warmed end-to-end policy latency from 198.2 ms in PyTorch to 73.8 ms on an NVIDIA GeForce RTX 5090 with CUDA 12.9 and TensorRT 10.16.1”
Source→“most VLA policies are still trained mainly as direct observation-and-language-to-action mappings, so their supervision for how the physical scene evolves under robot intervention is indirect compared with video-based world modeling”
Source→“DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.