Changwen Zheng
Changwen Zheng is a professor at the Institute of Software, Chinese Academy of Sciences, where he also serves as deputy director of the National Key Laboratory of Integrated Information System Technology. His research spans machine learning, computer simulation, and information processing, with recent work focusing on reinforcement learning post-training methods for large language models and vision-language-action policies. He is known for contributions including evolutionary route planning for unmanned aerial vehicles and information-theoretic reinforcement fine-tuning frameworks for LLMs.
“Temporal GRPO keeps changes on the preceding stages close to zero and produces the largest positive improvement at md, indicating that it preserves acquired preceding behaviors and concentrates the update on the stage responsible for the rollout difference.”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.