#machine learning
Feb 2026
Constrained Group Relative Policy Optimization
This work introduces Constrained GRPO, a Lagrangian-based extension of GRPO for constrained policy optimization, and addresses the coupling induced by reward scalarization by scalarizing standardized advantages rather than rewards.
Roger Girgis, Rodrigue de Schaetzen, Luke Rowe et al.
· arXiv.org · 2 citations
· ⚡1