Jul 2026
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
This work develops LEMUR: Learning to Align with Multi-Objective Reinforcement Learning with Preference feedback, a novel framework where an agent interactively learns from the preferences of multiple humans to learn optimal multi-objective policies.
Manith Adikari, Bei Peng, Samuele Vinanzi et al.
· arXiv.org · 0 citations