Generate, Track, Improve: Perceptive Multi-Skill Humanoid Locomotion with RL-Fine-Tuned Motion Generators
1. Key Themes
Off-Policy RL Fine-Tuning of Generative Motion Models is Dramatically More Sample-Efficient Than On-Policy Alternatives
The paper's central contribution is an off-policy RL fine-tuning loop using Advantage Weighted Regression (AWR) to improve a flow matching motion generator. The key finding is that this approach is orders of magnitude more sample-efficient than the obvious alternative (PPO on a residual policy). The authors state: "we stopped training after 2,000 iterations at which point it had used more than 160x the data of the off policy algorithm and took more than 30x the wall clock time. At this point, the PPO baseline achieves a success rate around 13 percentage points below the AWR's" (Section III.B). For anyone building humanoid control systems, this means the cost of adapting generative locomotion policies to new environments can be dramatically reduced by choosing the right RL algorithm — AWR over PPO.
Raw Depth Images Eliminate the Odometry and Height-Map Dependency That Plagues Deployment
Most perceptive humanoid locomotion systems require height maps and odometry, which create significant deployment friction (requiring height scanning packages, accurate state estimation, and careful calibration). This paper demonstrates that conditioning both the generator and tracker directly on raw depth images from two cameras removes this dependency entirely. The authors note: "By using raw depth images to perceive the environment no odometry or height maps are needed, and outdoor deployment is easy" (Abstract) and "By using the raw depth camera readings the transfer to the outdoors is easy and no adjustments are needed relative to the inside" (Section III.D). This is a deployment simplification that directly reduces engineering overhead for real-world humanoid operations.
Two-Camera Architecture Enables Autonomous Speed Modulation for Terrain Traversal
A single downward-facing camera — the standard in many perceptive locomotion systems — is insufficient for dynamic locomotion at speeds up to 2.5 m/s because the robot cannot see terrain early enough to adjust its approach. The paper demonstrates that adding a forward-facing upper camera is critical: "Removing the upper camera hurts the ability to traverse a given terrain by up to 43% as demonstrated with the box terrain" (Table IV, Section III.F). The upper camera enables the policy to "see terrain coming from further away and adjust its velocity regardless of the commanded speed so it can traverse the terrain" (Abstract). This means the velocity command can be decoupled from terrain awareness — the robot autonomously slows down when approaching obstacles and resumes speed after clearing them (Figure 6, Section III.C).
Layered Generator-Tracker Architecture Enables Modular Skill Scaling
The architecture separates motion planning (generator, queried every 0.24s) from motion execution (tracker, running at 50 Hz), creating a natural computation split. The authors argue this is "ideal for scaling up generality in humanoids" because "the generator and tracker paradigm leads to modularity through a layered architecture and the ability to easily add additional skills, provided a general enough tracking controller" (Section I.A). A single policy pair handles walking, running, standing, jumping on/off boxes, and stair traversal — and the authors claim it "can be easily extended to many more skills" (Section IV).
2. Contrarian Perspectives
Fine-Tuning the Generator, Not the Tracker, Is the Key Lever for Adapting to New Environments
Most robotics teams facing locomotion failures on new terrain would fine-tune the tracking controller — the low-level policy that actually controls the robot. This paper argues the opposite: you should fine-tune the high-level motion generator instead. The authors explicitly state: "Adjusting this choice of mode is not something that can be achieved with fine tuning the tracker policy, making this a complementary tool" (Section III.A). The generator decides what motion to produce (e.g., stairs gait vs. box jump), and if it picks the wrong mode, no amount of tracker improvement can fix that. The 80-percentage-point improvement in skill selection (Figure 4c) is entirely attributable to generator fine-tuning, something tracker fine-tuning fundamentally cannot address.
Generative Models for Locomotion Can Be RL-Fine-Tuned Without Adding Gaussian Noise to Actions
The conventional approach to RL fine-tuning of generative policies (diffusion or flow matching) is to add Gaussian noise to the action outputs for exploration, as done in DPPO. This paper argues that for high-dimensional action spaces (their generator outputs ~3,000 values per plan), this approach fails: "independent gaussian noise on those actions does not lead to the structured trajectories that would create better references" (Section I.B). Instead, they perturb the conditioning (velocity commands, initial noise samples) and spawn the robot in varied positions, allowing structured exploration that produces reasonable trajectories. This challenges the assumption that standard PPO-style exploration generalizes to high-dimensional generative locomotion policies.
3. Companies Identified
Unitree, Robotics hardware manufacturer, Their G1 humanoid is the deployment platform for all hardware experiments. "A single policy pair enables a Unitree G1 humanoid to walk, run, stand, jump on and off of boxes, and traverse stairs in outdoor environments" (Abstract). The G1's 29-DOF body is the target system, making this work directly relevant to anyone deploying Unitree G1 robots.
NVIDIA, GPU and edge computing company, Their Jetson Thor runs the deployed policies and Jetson Orin handles depth processing. "To deploy the policies on the hardware we choose to run them on an NVIDIA Jetson Thor. Flow matching inference time is around 11 ms while the tracker is less than 1 ms" (Section II.F). IsaacLab is also used for simulation training. NVIDIA's hardware stack is the deployment target, and inference latency numbers (11ms for generator, <1ms for tracker) are practical benchmarks for system designers.
Stereolabs (ZED), Depth camera manufacturer, Their Zed X and Zed X mini cameras provide the depth perception. "The two cameras are a Zed X mini and Zed X which both feed into a NVIDIA Jetson Orin for depth processing and down sampling before being sent over ROS2 to the Thor" (Section II.F). The specific camera choice and placement (upper forward-facing + lower downward) is a deployment detail others can replicate.
Physical Intelligence, VLA company, Referenced in related works for RL fine-tuning of flow-based VLAs. Cited as [39] for π*0.6, showing the connection between RL fine-tuning of generative models in manipulation (VLAs) and this work's application to locomotion (Section I.A).
4. People Identified
Zachary Olkin, Caltech (Department of Control and Dynamical Systems), Lead author and presumably lead implementer. His prior work on "Chasing Autonomy" (reference [4]) established the CLF-RL running framework that this paper builds upon. This work extends his trajectory from dynamic running to full multi-skill perceptive locomotion.
Aaron D. Ames, Caltech (Department of Control and Dynamical Systems), Senior author and PI. Ames is a leading figure in control-based humanoid robotics, known for CLF (Control Lyapunov Function) methods. His lab's approach blends classical control theory with modern RL, which is reflected in this paper's use of CLF-RL for the tracker and the structured, theory-informed RL fine-tuning of the generator. The work is supported by the Technology Innovation Institute (TII).
William D. Compton, Caltech, Co-author who also contributed to the related "Terrain Consistent Reference-Guided RL for Humanoid Navigation Autonomy" (reference [10]), which uses raw depth scans with self-scanning. This paper builds on that perceptive locomotion foundation.
5. Operating Insights
The Generator-Tracker Split Maps Naturally to Edge Compute Constraints
The architecture exploits a real-world compute constraint: high-frequency control (50 Hz) requires a lightweight policy, while motion planning can run at lower frequency (every 0.24s = ~4 Hz). The tracker is a small MLP with hidden dimensions [1024, 512, 256, 128] running in <1ms, while the generator is a 26.8M-parameter transformer running in ~11ms. For teams designing humanoid compute stacks, this validates a two-tier inference architecture where a small fast policy handles control and a larger slower policy handles planning — and both fit on a single Jetson Thor. The transformer's prefix encoder (depth image + velocity conditioning) is computed once and cached across all 8 flow matching integration steps, further reducing inference cost (Section II.D).
Data Pipeline Quality Directly Caps Final Policy Performance
The authors are explicit that the quality of the motion library used for training bounds the final system's velocity tracking: "these motion clips upper bound the velocity tracking performance of the final policies since they are trained to track these clips. Therefore any lost tracking capabilities here can not be recovered with the current pipeline" (Section III.E). Their optimized references achieve statistically significant improvements over Motion Bricks alone (Table III), particularly at higher speeds (0.442 vs 0.896 m/s RMSE for velocities 1.0-2.5 m/s). For teams building locomotion systems, this means investing in the data pipeline (motion optimization, terrain retargeting) pays off more than tuning the RL algorithm — a bad reference library cannot be fixed downstream.
6. Overlooked Insights
RL Fine-Tuning Can Incorporate Pre-Training Data to Prevent Catastrophic Forgetting — At Near-Zero Cost
A practical detail buried in Section II.E: during RL fine-tuning, the authors mix in original pre-training data with weight 1 (no advantage weighting) into the flow matching loss update, while excluding it from critic learning and advantage labeling. This simple technique — "we can easily mitigate any potential issues with forgetting previous behaviors" — means you can adapt to new terrain without losing existing skills, and it requires no additional architectural complexity. For teams iterating on deployed humanoid policies, this is a lightweight recipe for continual learning that avoids the common pitfall of catastrophic forgetting when fine-tuning on new environments.
The Policy Generalizes to Outdoor Environments Without Simulating Them
The authors note that the hardware deployment worked outdoors "with walls and trees and bushes without needing to simulate every possible perceptive condition" (Section III.D). The training simulation included terrain geometries but not the full visual complexity of outdoor scenes. The depth-only perception pipeline (no semantic information, no RGB) combined with domain randomization on camera angles/offsets and depth noise was sufficient for sim-to-real transfer to unstructured outdoor environments. This suggests that for locomotion (as opposed to manipulation), depth-only perception with aggressive domain randomization may be more robust than more complex perceptual architectures — and far cheaper to simulate.