Developers of the Qwen3-VL-2B vision-language model used as the backbone of GTA-VLA.
“Qwen3-VL-4B (Bai et al. 2025) introduced the concept of treating bounding box coordinates as tokens, allowing for explicit grounding”
Source→“CSD provides explicit task-conditioned spatial supervision before action learning... The teacher is used exclusively to estimate the target-object center from the original RGB observation and task instruction... Qwen3-VL-4B and explicit coarse spatial labels are not used during downstream policy training or deployment”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.