RSPO: Regularized Self-Play Alignment of Large Language Models
It is shown that RSPO with appropriate regularizers can substantially improve the length-controlled win rate on AlpacaEval-2 across a range of base models, while also achieving consistently superior performance on Arena-Hard, MT-Bench, ArmoRM, and response diversity.
Xiaohang Tang, Sangwoong Yoon, Seongho Son et al.
· 6 citations
· ⚡1