MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models
The proposed Module Level Reward Evolution Framework integrates three mechanisms: reflection-based refinement, hybrid credit assignment, and a merge strategy with rollback, which together improve the effectiveness and robustness of reward optimization.
Chenglin Liu, Xun Wang, Ruishuo Chen et al.
· 0 citations