Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
1. Key Themes
Autonomous Skill Refinement Without Weight Updates
RPG improves robot capabilities by refining the code (skills and system prompts) rather than fine-tuning neural network weights. This is a massive shift for physical AI deployment, as it avoids the massive compute and data requirements of traditional model training. As the paper states: "We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights." (Abstract). By treating skills as a persistent, editable library, robots can get smarter over time through software updates rather than hardware or model upgrades.
Sim-to-Real Transfer via Frozen Systems
The system practices autonomously in simulation, gets frozen, and deploys to the real world with only basic physical calibration. This proves that the skills learned in simulation are robust enough for physical reality without requiring continuous online learning on the robot. The authors note: "After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks." (Abstract). For operators, this means you can develop and validate complex manipulation logic safely and cheaply in sim before pushing it to a fleet of robots.
Cross-Task Generalization of Skills
Improving a skill for one task naturally benefits other tasks that rely on the same fundamental manipulation primitives. This creates a compounding return on investment for engineering effort. The paper highlights: "RPG improves both the shared skill library and the system prompt that guides a multimodal agent in using it. Changes to these shared components persist across episodes and can benefit other tasks that reuse them." (Sec. I). Furthermore, the authors show that tasks that received no direct improvement feedback still improved because the underlying shared skills were repaired (Appendix D, Table VII).
2. Contrarian Perspectives
Better Skills Beat Better Models
The prevailing wisdom in AI is that scaling up model size and capability is the primary driver of performance. This paper argues that for robotics, improving the execution system (the code and prompts the model uses) yields better results than simply buying a more powerful LLM. The authors state: "These results suggest that improving the shared execution system can provide larger gains than replacing the runtime model alone." (Sec. IV-E). Specifically, RPG using Gemini 3.8 Flash achieved a 95.0% success rate, vastly outperforming a baseline using the more powerful GPT-6 Astra Pro, which only achieved 60.0% (Table II).
Simulation Practice Doesn't Need to Perfectly Match Reality
Many robotics companies obsess over high-fidelity sim-to-real matching, attempting to perfectly reconstruct real-world scenes in simulation. RPG argues that you only need analogous tasks to practice the underlying manipulation capability. The paper explains: "The practice task may cover part of the source task or an analogous manipulation, rather than reproduce the source task or its exact scene geometry." (Sec. III-B). For example, practicing "organizing sunglasses" in the real dataset was translated to "closing a case that starts open and loaded" in simulation (Table I). This drastically lowers the barrier for generating useful simulation training environments.
3. Companies Identified
Amazon
Description: E-commerce and cloud computing giant. Why relevant: Several authors, including Pieter Abbeel and Haozhi Qi, are affiliated with Amazon FAR (Foundational AI Research), and lead author Yen-Jen Wang was an intern there. This indicates Amazon's continued investment in foundational robotics and embodied AI research. Quotes: "2Amazon FAR" (Author affiliations); "Yen-Jen Wang and Haoru Xue were interns at Amazon FAR during this work." (Footnote).
Google (DeepMind)
Description: AI research and cloud computing leader. Why relevant: Google's Gemini 3.8 Flash model is used as the core runtime agent for online decision-making and task-level code generation in the most successful RPG configuration. Quotes: "RPG, ASPIRE, and RATs use Gemini 3.8 Flash for online decision making and task-level code generation." (Sec. IV-B).
OpenAI
Description: AI research and deployment company. Why relevant: OpenAI's GPT-6 Astra Pro was used as the strongest baseline model for the CaP-Agent0 configuration, which RPG significantly outperformed. OpenAI models were also used for manuscript editing. Quotes: "outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%)." (Abstract); "OpenAI ChatGPT and Anthropic Claude were used to assist with language editing and manuscript polishing." (Sec. VI).
4. People Identified
Pieter Abbeel
Lab/Institution: UC Berkeley / Amazon FAR Why notable: A leading figure in robot learning and reinforcement learning. His involvement signals that this approach to skill libraries and LLM-driven control is gaining traction among top-tier academic and industry researchers. Quotes: "Pieter Abbeel†,1,2" (Author list).
Yen-Jen Wang
Lab/Institution: UC Berkeley / Amazon FAR Why notable: Lead author of the paper, driving the research on autonomous self-improvement for embodied agents. Quotes: "Yen-Jen Wang∗,‡,1" (Author list).
Haozhi Qi
Lab/Institution: University of Chicago / Amazon FAR Why notable: Equal advising author alongside Pieter Abbeel, indicating a key role in directing the research strategy. Quotes: "Haozhi Qi†,2,4" (Author list).
S. Shankar Sastry
Lab/Institution: UC Berkeley Why notable: A renowned control theorist and roboticist, bringing deep theoretical rigor to the practical deployment of these systems. Quotes: "S. Shankar Sastry1" (Author list).
5. Operating Insights
Treat Skills as Code, Not Just Weights
CTOs should focus on building robust, verifiable skill libraries that LLMs can call, rather than trying to train end-to-end neural policies for every new task. This allows for human-readable debugging, modular updates, and cross-task generalization. The paper argues: "A scalable system should combine these strengths by delegating repeatable execution to reusable code and reserving multimodal reasoning for decisions that benefit from visual or semantic judgment." (Sec. I). Over 15 rounds, the skill library grew from 15 to 38 entries, with 66 modifications to existing skills, all without touching model weights (Sec. IV-B).
Use Privileged Simulation State for Diagnosis, Not Execution
You can use ground-truth simulation data to figure out why a robot failed, and then fix the code, without letting the deployed robot cheat by relying on that ground-truth data. RPG runs a "Privileged Agent" alongside the "Runtime Agent" to compare executions. The authors explain: "RPG instead uses privileged state in a separate execution to support failure diagnosis... Its successful executions can guide improvements to the Runtime Agent, while failures shared by both agents can expose defects in the shared skills." (Sec. III-C). This is a highly efficient way to generate targeted training signal for code revision.
6. Overlooked Insights
System Prompt Engineering is a Major Bottleneck for Agents
While the paper focuses on skill libraries, the single largest jump in performance came from fixing the system prompt to stop the robot from wasting its observation budget. The authors note: "The largest improvement occurs between rounds 3 and 4, when success increases from 43.2% to 73.6%. This round changes the system prompt but records no skill edit. The new prompt introduces a bounded perception–action loop and tracks the remaining interaction budget, encouraging the agent to act on its current estimates instead of repeatedly requesting observations." (Sec. IV-B). For anyone deploying LLM-based agents, managing the interaction budget via prompt engineering is critical to preventing timeouts and failures.
Deformable Object Manipulation Remains a Hard Boundary
While RPG excels at rigid body manipulation by adding procedural checks (e.g., test lifts, path validation), it struggles significantly with cloth because the state changes too dynamically for simple code checks to verify. The authors admit: "For deformable objects, the underlying state itself changes substantially under contact, and visually similar configurations can correspond to different layering and attachment states. Improving this regime likely requires richer representations of deformable state and contact, rather than only additional composition of the existing rigid-object skills." (Appendix F). Investors should be wary of companies claiming general manipulation capabilities if their core technology relies solely on rigid-body procedural logic without specialized representations for deformables.