Jul 2026
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
This work proposes PRISM, a new multi-reward RL framework built upon the idea of policy-space decomposition and composition, which alleviates the potential conflict during multi-reward policy optimization, while enabling controllability during inference by flexible policy composition.
Ruiming Liang, Yin-Jie Zhong, Yizhen Yuan et al.
· arXiv.org · 1 citation