InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
1. Key Themes
Test-Time Skill Acquisition Without Retraining
InterEvolve's core contribution is enabling a humanoid robot to solve tasks it was never trained on — without retraining the controller. The system repurposes existing motor competence by evolving "reward programs" (staged reward functions with tunable constants) at test time. An LLM agent iteratively revises these programs based on execution feedback from parallel simulation, while a numerical optimizer (CMA-ES) tunes the constants. The paper states: "We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining" (Abstract). This is significant because it shifts the cost of adapting to new tasks from expensive RL training runs to inference-time compute in simulation.
Object-Aware Behavioral Foundation Model
The paper introduces an extension to forward-backward (FB) behavioral foundation models that makes them object-aware. Prior FB models (like BFM-Zero) observe only the body, so "rewards that differ only in the object's goal can collapse into one behavior" (Sec. 2). InterEvolve adds trainable object residuals to frozen body networks: "we attach trainable object residuals that read object features to its frozen body networks, and we train these residuals on large-scale human-object interaction (HOI) data" (Sec. 1). This is what enables a single fixed controller to execute diverse object-interaction behaviors specified by reward programs. The object-aware model achieves 60% tracking success vs. 8% for body-only BFM-Zero (Table 1).
LLM-Driven Reward Program Evolution
The system uses an LLM agent (DeepSeek-V4-Flash) to write and iteratively revise reward programs — structured code specifying staged rewards, completion conditions, and tunable constants. The agent receives execution feedback (per-criterion pass rates, stage traces, failure diagnostics) and proposes structural revisions, while CMA-ES handles constant tuning. The paper shows this matters enormously: "InterEvolve more than doubles the success of the best calibrated program" reaching 86.5% success vs. 34.6% for a calibrated agent-written program without evolution (Table 2). The gain comes specifically from structural changes — "what is rewarded and how objectives are staged" — not just better hyperparameters.
Skill Library Accumulation and Composition
Verified reward programs are stored in a text-based skill library that the LLM agent can reference for future tasks. This enables long-horizon composition: "a later long-horizon composite task, such as carrying a box, placing it, and then kicking it, or a novel box tip, can build on earlier experience" (Sec. 1). The full library solves 8/10 relocation tasks, 4/10 stacking tasks, and 5/10 carry-place-kick tasks, vs. nearly zero without any library (Table 5). Critically, reuse depends on covering the right contact modes: "Reuse therefore depends on covering the contact modes a task needs, not on having any stored experience" (Sec. 4.4).
Real-World Deployment on Unitree G1
The system was deployed on a physical Unitree G1 robot using only egocentric onboard perception. "The robot detects the box with its egocentric camera, and FoundationPose estimates its 6-D pose, which supplies the controller's object features. Each program is evolved and verified in simulation, then run on the robot" (Sec. 4.3, Figure 4). This demonstrates a sim-to-real pipeline where reward programs evolved in simulation transfer directly to hardware.
2. Contrarian Perspectives
Human-Designed Rewards Are Fundamentally Inadequate for Loco-Manipulation
The paper argues that expert reward engineering leaves most of a controller's capability untapped. Even a human-designed reward calibrated with CMA-ES achieves only 18% success across eight task families, vs. 86.5% for evolved programs (Table 2). The paper states: "a naively handcrafted reward underperforms even if the controller contains the relevant motor capabilities" (Sec. 1). This challenges the common practice in robotics companies of relying on skilled ML engineers to hand-design reward functions — the paper shows that structural reward design (which stages to use, what to reward) matters more than parameter tuning, and that automated search over program structure discovers strategies humans miss. Figure 1 shows a concrete example: "A human-designed reward tries to hack this behavior but fails. The reward program that InterEvolve evolves compensates with novel body used to succeed."
You Don't Need to Retrain for Every New Task — Your Controller Already Has the Skills
The central thesis challenges the dominant paradigm of training task-specific policies. The paper argues: "a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control" (Abstract). The evidence in Figure 11 is striking: evolved behaviors like lifting onto a table, chest-height carrying, kicking, and tipping "form their own clusters" distinct from training data, meaning the system produces motions that "no training reference contains" (Sec. D.7). This suggests companies investing heavily in per-task data collection and policy training may be over-investing — the bottleneck is accessing existing competence, not adding new training.
Simulation Compute, Not LLM Compute, Is the Bottleneck for Test-Time Adaptation
While the AI industry is focused on LLM inference costs, the paper reveals that for physical AI, simulation dominates: "Each round spends 1–2 min in the language model and 16–34 min in simulation on one GPU, so simulation dominates wall-clock time" (Sec. C.4). The LLM costs are trivial (0.23M tokens for a full 5-round evolution). This implies that faster simulators (the paper explicitly mentions mjlab as an alternative to Isaac Lab) would have more impact on adaptation speed than better language models.
3. Companies Identified
Unitree Robotics
- Description: Manufacturer of the G1 humanoid robot
- Why relevant: The physical deployment platform for InterEvolve. The controller is retargeted to the G1 with both rubber hands and Inspire dexterous hands. All real-world results use this platform.
- Quote: "evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception" (Abstract)
Inspire Robots
- Description: Manufacturer of dexterous hands used on the G1
- Why relevant: The paper demonstrates that InterEvolve extends to dexterous whole-body manipulation using Inspire hands on the G1, showing the approach generalizes beyond simple grippers.
- Quote: "We retarget both to the Unitree G1 (Unitree Robotics) with rubber hands and to the G1 with Inspire hands (Inspire Robots)" (Sec. 4.1)
DeepSeek
- Description: AI company providing the LLM used for reward program generation
- Why relevant: DeepSeek-V4-Flash is the LLM agent that writes and revises reward programs. The choice of a relatively lightweight model (Flash variant) suggests this task doesn't require frontier-scale reasoning.
- Quote: "DeepSeek-V4-Flash (Xu et al., 2026a) writes the reward programs" (Sec. 4.1)
NVIDIA (implied)
- Description: Creator of Isaac Lab simulation framework
- Why relevant: Isaac Lab is the simulation environment used for all training and evolution. The paper notes simulation is the cost bottleneck and suggests alternatives like mjlab could help.
- Quote: "All controllers run in Isaac Lab (Mittal et al., 2025)" (Sec. 4.1); "a faster simulator such as mjlab (Zakka et al., 2026) in place of Isaac Lab (Mittal et al., 2025) could shorten evolution substantially" (Sec. 4.3)
4. People Identified
Zhuo Lin (co-first author)
- Lab/Institution: University of Illinois Urbana-Champaign
- Why notable: Co-developed the InterEvolve framework. Equal contribution with Sirui Xu.
Sirui Xu (co-first author)
- Lab/Institution: University of Illinois Urbana-Champaign
- Why notable: Co-first author with prior work on human-object interaction (InterMimic, InterPrior). Has a track record in physics-based interaction control, which is the foundation for this work.
- Quote: Co-developed the object-aware FB model and the evolution framework.
Yu-Xiong Wang (co-advisor)
- Lab/Institution: University of Illinois Urbana-Champaign
- Why notable: Advisor with expertise in computer vision and learning. Co-advised the work.
Liang-Yan Gui (co-advisor)
- Lab/Institution: University of Illinois Urbana-Champaign
- Why notable: Co-advisor. The lab has produced a series of related works on human-object interaction (InterMimic, InterPrior, ULTRA), suggesting a coherent research program around humanoid loco-manipulation.
Yitang Li (referenced, not author)
- Lab/Institution: Carnegie Mellon University (with Kris Kitani, Guanya Shi et al.)
- Why notable: First author of BFM-Zero, the behavioral foundation model that InterEvolve builds upon. BFM-Zero is the body-only FB model that InterEvolve extends with object awareness.
Tairan He (referenced)
- Lab/Institution: Carnegie Mellon University
- Why notable: Key figure in humanoid teleoperation and control (OmniH2O, SONIC). His work on human-to-humanoid whole-body control is part of the foundation InterEvolve builds on.
5. Operating Insights
The Controller-Program Separation Changes How You Should Architect Robot Software
InterEvolve demonstrates a clean separation: a frozen motor controller handles low-level execution, while task knowledge lives in inspectable, editable reward programs. This means you can update task behavior without touching the controller — no retraining, no fine-tuning, no data collection. The paper states: "A behavioral foundation model learns how to move once, and task knowledge lives outside its weights, in reward programs that a language model can read, edit, and test by execution" (Sec. 5). For a CTO, this means: invest in a strong general controller once, then iterate on task logic through program search. This is architecturally similar to how LLMs separate foundation model from prompting — and it brings the same benefits of rapid iteration and inspectability.
Multi-Stage Decomposition Is Non-Negotiable for Contact-Rich Manipulation
The ablation in Table 3 shows that forcing single-stage programs causes the largest performance drop (from 86.5% to 44.7% success), more than removing CMA-ES tuning (51.6%) or multi-scenario evaluation (68.4%). The paper explains: "contact-rich interaction needs a task decomposed into phases with their own objectives" (Sec. 4.3). For engineering teams building manipulation systems, this means task decomposition (acquisition → transport → placement → release) should be a first-class architectural concern, not an afterthought. The reward program representation makes stages explicit and independently editable, which is what enables the LLM agent to effectively search over strategies.
Expect ~2 GPU-Hours and 5 Rounds for New Task Adaptation
The cost accounting in Sec. C.4 provides concrete numbers: a five-round evolution on one task family executes ~9.5M environment steps on average, costing about 2.1 GPU-hours plus 0.23M LLM tokens. Each round takes 16-34 minutes of simulation. This is fast enough for practical iteration — a team could adapt to a new task in under 2 hours on a single GPU. However, this assumes the controller already has the relevant motor competence; tasks requiring truly novel motor skills (like kicking, which is "rare in the training data") achieve lower success (59.4%) even after evolution.
6. Overlooked Insights
The Reward-Inference Bank Is a Critical and Fragile Component
The paper reveals that the "bank" — a fixed set of ~50K states used to convert rewards into latent prompts — has a goldilocks problem. Too small (6.25K states) and success drops to 21.9%; too large (200K states) and it drops to 14.1% because "its tilted weight concentrates on an order of magnitude fewer states than the 100k bank" (Table 9, Sec. B.3). The reward tilt parameter β similarly has a sweet spot: β=10 gives 59.4% success while β=30 drops to 28.1% because "nearly all weight [goes] on three states." This means the bank construction and tilt calibration are critical engineering decisions that could make or break deployment, yet they're buried in the appendix. A team deploying this approach would need to carefully tune these parameters per task category.
Programs Don't Transfer Across Objects Without Adaptation
Table 14 shows that reward programs evolved for the large box transfer poorly to other objects: averaged across families, unchanged programs reach only 36-59% success on new boxes. Push-to-mark drops from 100% to 0% on the plastic and small boxes. However, brief adaptation (re-running the search for the new object) raises averages to 79-89%. This means the "no retraining" claim has an important caveat: while you don't retrain the controller, you do need to re-run the evolution loop for each new object geometry. The skill library helps (programs provide good starting points), but direct transfer is not reliable. This has implications for deployment in environments with diverse object types — the system needs time to adapt per object category.