Aileen Liao
Aileen Liao is a PhD student in Computer and Information Science at the University of Pennsylvania GRASP Lab, co-advised by Michael Posa and Dinesh Jayaraman. She works on learned controllers and embodied intelligence in robotics. She is best known as a co-author of LENS (LLM-guided Environment Simplification), a method that uses vision-language models to generate task-relevant scene abstractions for robotic manipulation in cluttered environments.
“On hardware, the paper shows cumulative success rates reaching 80% across three feedback iterations, with successes distributed across all three iterations rather than concentrated in the first (Figure 6, Section 5.2).”
Source→“LENS uses a vision-language model (GPT-4o) not to generate actions or plans, but to simplify the scene before any planner or controller runs.”
Source→“LENS requires no task-specific engineering, no model fine-tuning, and no data collection. It uses off-the-shelf GPT-4o with prompt engineering. The VLM query time averages 1.76 seconds (Section 5), negligible relative to execution time.”
Source→“The paper demonstrates this works as a front-end across three fundamentally different paradigms — Task and Motion Planning (TAMP), contact-implicit model predictive control (C3+), and the π0.5 vision-language-action model.”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.