#artificial intelligence
Feb 2026
Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment
Personalized GRPO is introduced, a novel alignment framework that decouples advantage estimation from immediate batch statistics and achieves faster convergence and higher rewards than standard GRPO, thereby enhancing its ability to recover and align with heterogeneous preference signals.
Jialu Wang, Heinrich Peters, A. Butt et al.
· arXiv.org · 1 citation