Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/ARXIV PHYSICAL AI RESEARCH/CoDance: Learning Reactive and C…
PAPR
// RESEARCH PAPER
ARXIV PHYSICAL AI RESEARCH

CoDance: Learning Reactive and Compliant Human-Humanoid Interaction from Video

DATE October 4, 2026SOURCE ARXIV PHYSICAL AI RESEARCHPARTICIPANTS ZHUOQUN CHEN, SHUCHENG JIA, BOYUAN CHENARXIV 2610.05324
// SUMMARY

1. Key Themes

Learning Complex Interaction from a Single Video

CoDance demonstrates that a humanoid robot can learn a complex, physically interactive behavior—partnered dancing—using only a single 9.7-second video of two human dancers. The system reconstructs the 3D motions of both individuals, retargets one to the robot and the other to a simulated partner, and trains a policy that coordinates footsteps and maintains two-hand contact. As stated in the Abstract: "Given a single video of two human dancers, CoDance retargets their motions into a robot reference and a moving partner." This drastically lowers the data collection barrier for teaching robots contact-rich, interactive skills.

Generating Force-Aware Training Data from Kinematic Demonstrations

A major hurdle in learning physical interaction from video is that videos contain motion but no force data. CoDance solves this with a "multi-link compliance augmentation" that simulates structured interaction forces at the robot's hands and adapts the reference motion accordingly. The paper notes in Section I: "We address this with a multi-link compliance augmentation that transforms the kinematic demonstration into force-aware training data." This allows the robot to learn how to yield and comply to forces without requiring expensive force-capture suits or teleoperation during the data collection phase.

Reactive Partner Following Without Explicit Velocity Commands

Instead of receiving a joystick command or a pre-programmed trajectory, the robot observes sparse keypoints (pelvis and knees) of its partner and adjusts its locomotion in real-time. The paper highlights in Section III-B: "The policy does not take a velocity command. It follows the partner from the partner’s keypoints." In simulation, masking or reversing the partner's observed motion caused the robot to fail or step incorrectly, proving the policy genuinely reacts to the human rather than just replaying a memorized sequence (Table II).

Successful Hardware Deployment on a Humanoid

The framework was validated on a physical Unitree G1 humanoid, achieving sustained two-hand dancing with a human partner. The robot maintained a consistent following distance (mean of 0.48m to 0.50m against a 0.5m target) and successfully transitioned between forward and backward motions for up to 38.9 seconds without falling (Section IV-G). This shows the simulation-trained policy successfully transfers to the real world despite dynamics gaps.

2. Contrarian Perspectives

Deviations from Reference Motion Are Features, Not Bugs

Conventional motion imitation treats any deviation from a reference trajectory as a tracking error to be minimized. CoDance argues that for physical interaction, deviating from the reference is the correct behavior when responding to external forces. Section I states: "This creates a direct conflict with conventional motion imitation, where deviations from the reference are treated as tracking errors." By training on "adapted references" that encode compliant responses to simulated forces, CoDance teaches the robot that moving away from the original kinematic demonstration is necessary for safe interaction.

Random Pushes Are Unnecessary for Robustness

A standard practice in reinforcement learning for robotics is applying random external pushes to the robot during training to improve balance and robustness. CoDance explicitly disables this, arguing that their simulated interaction forces serve the same purpose more effectively. Section III-B notes: "Random pushes are disabled because the scheduled interaction events already perturb the robot." This suggests that for contact-rich tasks, domain randomization should be structured around the expected interaction physics rather than generic noise.

3. Companies Identified

Unitree

Description: Manufacturer of the G1 humanoid robot. Why relevant: The physical deployment and validation of the CoDance policy were performed on a Unitree G1. The paper states in Section III-C: "We deployed the converged decoupled policy on a Unitree G1 at 50 Hz. Joint position targets were sent through the Unitree SDK at 200 Hz."

Vicon

Description: Manufacturer of optical motion capture systems. Why relevant: Vicon systems were used to track the human partner's keypoints and the robot's pelvis during hardware deployment, bridging the gap between simulation and reality. Section III-C notes: "A Vicon system tracked the person’s pelvis and knees and the robot pelvis, providing the same partner keypoints used during training."

NVIDIA

Description: Manufacturer of GPUs used for training the reinforcement learning policies. Why relevant: The training pipeline relies on NVIDIA hardware for reasonable iteration times. Section III-B states: "Training converges in approximately 37 h for the decoupled policy and 29 h for the whole-body tracker on one NVIDIA L40S."

4. People Identified

Zhuoqun Chen, Shucheng Jia, Boyuan Chen

Lab/Institution: Duke University (General Robotics Lab). Why notable: The authors of the paper and developers of the CoDance framework. They successfully bridged the gap between video-based learning, force-aware data augmentation, and physical humanoid deployment. The paper notes: "All authors are with Duke University."

X. B. Peng

Lab/Institution: Referenced for prior work on AMP (Adversarial Motion Priors) and DeepMimic. Why notable: CoDance leverages Peng's AMP framework to learn the locomotion style of the lower body without explicitly tracking the reference, allowing the upper body to focus on compliance. Section III-B states: "The decoupled policy uses an adversarial motion reward [3]. Its discriminator compares two-frame transitions from policy rollouts with transitions from the training references using lower-body and torso poses in the pelvis frame."

G. Margolis

Lab/Institution: Referenced for SoftMimic. Why notable: CoDance's multi-link compliance augmentation builds directly upon the quasi-static compliance solve introduced in SoftMimic. Section III-A notes: "Our augmentation builds on the quasi-static compliance solve of SoftMimic [4]." Margolis's prior work laid the foundation for transforming kinematic references into force-aware data.

5. Operating Insights

Sparse Keypoints Are Sufficient for Reactive Locomotion

CTOs and engineering leads should note that full-body tracking of a human partner is not required for reactive robot locomotion. CoDance successfully trained the robot to follow a partner using only the partner's pelvis and knee positions observed over the last second. Section III-B states: "Partner observations contain the pelvis and both knees in the robot pelvis frame over the last second." This significantly simplifies the perception stack required for human-robot collaboration in the real world.

Compliance Can Be Commanded via Stiffness Parameters

The system allows operators to dynamically control how stiff or compliant the robot's arms are by passing a commanded stiffness value. During hardware deployment, the wrists were set to a stiffness of 140 N/m. Section III-C notes: "Both wrists used a commanded stiffness of 140 N m−1. Physical forces came directly from the person’s hands during the two-hand hold." This provides a tunable knob for safety and interaction quality without retraining the policy.

Clip Chaining Is Critical for Long-Horizon Tasks

If a robot is trained on short, isolated clips of behavior, it will likely fail when trying to string those behaviors together over a longer horizon. CoDance found that training without chaining clips together resulted in a 50% failure rate at the transition point, even though the robot was still successfully following the partner. Section IV-F states: "The policy trained without clip-chaining succeeded in about half of the trials. Its foot error before termination matched the other two, so it did not lose the partner before failing. It failed at the clip switch while still following."

6. Overlooked Insights

Fixed Following Offset Limits Spatial Generalization

While the robot successfully follows a partner, the current implementation is constrained to a single, fixed relative position. The foot-following target keeps the robot at a strict 0.5m offset along a single direction. Section V notes: "The foot-following target also keeps the robot at a fixed offset along a single direction from the partner, so dances that change the partners’ relative placement are not covered." This means the current policy cannot handle complex spatial maneuvers like circling or side-by-side walking without retraining.

Policies Remain Highly Dependent on the Reference Clip

Despite being reactive to the partner's observed motion, the policy still heavily relies on the original kinematic reference to know what to do. If the partner moves in a way that completely contradicts the reference clip (e.g., moving backward when the reference expects forward motion), the robot struggles. Section V states: "The policies still depend on their reference clip, so only about a quarter of the trials finish a clip when the partner moves opposite to it." This indicates the learned behavior is more of a "reference-guided reaction" than a fully general partner-following capability.