Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/ARXIV PHYSICAL AI RESEARCH/VioLA: Learning Generalist Human…
PAPR
// RESEARCH PAPER
ARXIV PHYSICAL AI RESEARCH

VioLA: Learning Generalist Humanoid Control Policies from Human Data

DATE October 8, 2026SOURCE ARXIV PHYSICAL AI RESEARCHPARTICIPANTS MERT ALBABA, MARTIN RIEDMILLER, ET AL. (ARXIV PHYSICAL AI)ARXIV 2610.12435
// SUMMARY

1. Key Themes

Human Demonstrations as Robot Action Supervision via Shared Latent Spaces

VioLA's core innovation is using pretrained motion encoders that map both human and robot motion into the same latent space, allowing human video recordings to directly supervise a robot policy's action outputs. The training pool contains 140.6 million frames across 781 hours, of which 93.2% comes from human recordings (Table 1, Appendix A.1). This means the data bottleneck for humanoid training shifts from expensive teleoperation to existing human video datasets. As the paper states: "A human recording is therefore labeled in the policy's action space, and the training demonstration pool contains 140.6 million frames, 93.2% of them human" (Abstract).

Zero-Shot Generalist Humanoid Control Without Task-Specific Fine-Tuning

VioLA achieves 100% locomotion success (30/30 trials) and 88.6% manipulation success (31/35 trials) on a real Unitree G1 without any task-specific fine-tuning, across tasks including walking, bending, sitting, closing a laptop, hanging a coat, and rotating a chair. This dramatically outperforms released baselines: "GR00T N1.7 and Ψ0 reach 16.7% and 0%, respectively" for locomotion (Abstract). For manipulation, GR00T N1.7 completes 0/35 trials and Ψ0 completes 3/35 (Section 4.3). The key implication: existing humanoid generalist policies that require per-task teleoperated fine-tuning are being leapfrogged by an approach that learns from abundant human data.

Hierarchical Decomposition: Policy Predicts Motions, Controllers Handle Joints

VioLA separates "what motion to make" from "how to move joints." The generalist policy predicts 128-dimensional motion latents (64 body + 64 hand), while frozen pretrained controllers (SONIC for body, a separate module for hands) convert these into coordinated joint commands at 50 Hz. This decomposition means the policy doesn't need to learn balance, coordination, or joint-level control — it learns task-level motion selection. The paper notes: "Complex movements can be represented with a few motion tokens. A complete chair-rotation execution uses an average of 60 distinct body tokens and 31 hand tokens" over ~10 seconds (Section 4.4, Figure 5).

Backbone-Agnostic Architecture

The approach works across two VLA backbones (GR00T N1.7, π0.5) and one world-action model (DiT4DiT), all using the same training data and budget. GR00T performs best overall (100% locomotion, 88.6% manipulation), DiT4DiT matches on locomotion but drops to 25.7% manipulation, and π0.5 achieves 66.7% locomotion and 45.7% manipulation (Section 4.6, Figure 7). This suggests the latent action space is the critical innovation, not the backbone choice.

Human-Only Training Produces Real-Robot Behavior

A VioLA variant trained exclusively on human demonstrations (zero robot data in policy training) achieves 40% locomotion success on the real robot — successfully walking, bending, sitting, and standing from a chair (Section 4.5, Figure 6). This validates that encoded human motion provides "usable action supervision for real-robot behavior" (Section 4.5).


2. Contrarian Perspectives

Human Data Can Be Better Than Robot Data for Locomotion

Most robotics companies assume robot demonstrations are inherently superior to human demonstrations for training robot policies. VioLA's data shows the opposite for locomotion: human-then-mixed training achieves 30/30 locomotion success vs. 24/30 (80%) for robot-only training at the same number of updates. The paper states: "Human-then-mixed training thus improves locomotion success by 20 percentage points at the same number of training updates" (Section 4.5). The robot-only policy fails all five bending trials, suggesting robot teleoperation datasets have coverage gaps that human data fills.

Joint-Level Action Prediction Is the Wrong Abstraction for Humanoids

The dominant approach in humanoid VLA models (GR00T, Ψ0, Helix 02) is to have the policy predict joint commands directly. VioLA argues this conflates two learning problems — coordination and task selection — making the policy harder to train and requiring more robot data. The paper states: "A generalist policy that predicts joint commands must learn how to coordinate these movements while also learning which movement the instruction requires. Whole-body controllers already solve the coordination half" (Section 1). The evidence is stark: released GR00T N1.7 and Ψ0, which predict joint-level actions, achieve 16.7% and 0% locomotion success zero-shot, while VioLA achieves 100%.

More Robot Data Does Not Always Help

The conventional wisdom in robotics is "collect more robot demonstrations." But VioLA shows that adding robot data to human data actually hurts manipulation performance: robot-only training achieves 94.3% manipulation success vs. 88.6% for human-then-mixed training (Section 4.5). The robot-only variant succeeds on all five shampoo trials but fails on four kettle trials, while the mixed variant shows the opposite pattern. The paper acknowledges: "Human-then-mixed training therefore improves locomotion by six successful trials but completes two fewer manipulation trials. Its benefit depends on the task" (Section 4.5).


3. Companies Identified

NVIDIA — Creator of GR00T N1.7, the backbone VioLA builds on and outperforms. GR00T N1.7 uses a Cosmos-Reason2 vision-language backbone with flow-matching action transformer (Appendix A.3). Released GR00T achieves only 16.7% locomotion and 0% manipulation success zero-shot, suggesting NVIDIA's current approach of predicting joint commands requires task-specific fine-tuning that VioLA eliminates. VioLA's results could reshape NVIDIA's humanoid strategy.

Physical Intelligence (π0.5) — VioLA tests their π0.5 backbone, which achieves 66.7% locomotion and 45.7% manipulation success with VioLA's latent action space (Section 4.6). π0.5 has higher prediction delay (208.6ms vs. 160.2ms for GR00T) and drops the coat in 4/5 trials. Relevant quote: "The π0.5 variant also closes the laptop and places the dumpling in every trial, but it drops the coat after grasping it in four trials" (Section 4.6).

Figure AI — Referenced for Helix 02, which "predicts whole-body joint targets that a controller tracks" (Section 2). This is the joint-level prediction approach VioLA argues against. Figure's competitive position may be weakened if latent-space policies prove more data-efficient.

Unitree Robotics — Hardware provider (Unitree G1 with Inspire hands) and dataset contributor (UnifoLM-WBT, 4.45M frames). Their robot is the evaluation platform throughout. Also referenced: "Unitree G1 uses five-fingered Inspire hands, with six actuated joints per hand" (Appendix A.5).

Lightwheel AI — Created EgoStandard/EgoSuite, the human recording dataset that supplies 131M frames (93.2% of all training data). This dataset is critical infrastructure for VioLA's approach. Referenced as: "human recordings from EgoSuite, derived from EgoStandard (Lightwheel AI, 2026)" (Section 4.1).


4. People Identified

Mert Albaba — Lead author, ETH Zürich / MPI-IS / Vesoma. Conducted work during internship at Vesoma. Co-lead on the project.

Michael J. Black — MPI-IS. Founder of SMPL/MANO body and hand models, which are foundational to VioLA's human motion representation. His lab's decades of work on body capture directly enables the human motion encoding pipeline.

Andreas Krause — ETH Zürich. Head of the Institute for Machine Learning, one of Europe's leading ML groups. His involvement signals serious academic backing.

Martin Riedmiller — Vesoma. Known for foundational work in reinforcement learning (including the DQN-related work at DeepMind). His presence at Vesoma suggests the company is pursuing humanoid control seriously.

Georg Martius — MPI-IS / University of Tübingen. Director at MPI-IS with focus on autonomous learning and embodied AI.

Wieland Brendel — MPI-IS / ELLIS Institute Tübingen. Notable for work on adversarial robustness and now physical AI.


5. Operating Insights

Latent Action Spaces Dramatically Reduce Data Collection Costs

For any company building humanoid policies, the data strategy matters more than the model architecture. VioLA demonstrates that 93% of training data can come from human video recordings rather than expensive robot teleoperation. The converted pool is 781 hours of data, of which 727.84 hours are human recordings (Table 1). If your data pipeline relies primarily on teleoperation, you are paying ~10-50x more per demonstration than necessary for locomotion behaviors. The key requirement is reliable human body and hand motion annotation (SMPL/MANO fitting), which is a solved problem at scale.

Frozen Controllers Enable Rapid Backbone Iteration

Because the motion encoders and decoders remain frozen during policy training, you can swap VLA/WAM backbones without retraining the control stack. VioLA tested three backbones with the same frozen modules and same training budget (Section 4.6). For a CTO, this means the low-level control investment (SONIC body controller, hand decoder) is a one-time cost that amortizes across multiple policy architectures. The prediction delay differences are significant though: GR00T at 160.2ms, π0.5 at 208.6ms, DiT4DiT at 248.5ms (Table 3) — at 50Hz control, these translate to 8-12 control steps of latency where the robot continues moving under the previous chunk.

Real-Time Chunking Is Essential for Smooth Execution

Without RTC (real-time chunking), the robot experiences sharp discontinuities when a new action chunk replaces the current one — even during simple tasks like standing still. The paper measured maximum step changes dropping from 0.332 to 0.064 for Stand Still and 0.455 to 0.070 for Raise Arm when RTC is enabled (Section 4.7, Figure 8). Any deployment of chunked VLA policies needs this or equivalent smoothing, or the robot will jerk at every replanning boundary.


6. Overlooked Insights

Grasping and Object Retention Remain the Hard Frontier

While VioLA achieves strong manipulation results, the failure modes are revealing. The bottle-carrying task fails 5/5 across all variants, with failures at "grasp timing, the transition from grasping to walking, and upright release" (Appendix A.7). The hand decoder "produces finger-position commands without contact-force feedback" (Section 5, Limitations). This means the current architecture cannot close the loop on grip force — a fundamental limitation for any task requiring compliant grasping. Companies deploying humanoids for object transport should not assume latent-space policies solve the grasping problem.

The Training Schedule Matters as Much as the Data

VioLA's human-then-mixed schedule (150k steps human-only, then 50k steps uniform across all five corpora) is not arbitrary. The paper notes: "the VioLA schedule assigns an expected 80% of training examples to human data, distinct from their 93.2% share of the available frames" (Appendix A.1). Over-sampling human data too aggressively hurts manipulation (robot-only achieves 94.3% vs. 88.6% mixed), while under-sampling it hurts locomotion (robot-only achieves 80% vs. 100% mixed). The optimal mixing ratio is task-dependent and not yet solved — an open operational question for anyone deploying this approach.