MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models
Published in arXiv preprint, arXiv:2608.18827, 2026
MLREF evolves a persistent module pool for reward design in RL, reusing effective components across iterations and outperforming strong baselines by 25.2% in locomotion and 6.6% in manipulation.
Recommended citation: Chenglin Liu, Xun Wang, Ruishuo Chen, Zhuoran Li, and Longbo Huang. (2026). "MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models." arXiv preprint arXiv:2608.18827.
Download Paper
