Skip to content

Author

Xiaohang Tang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Feb 2025

RSPO: Regularized Self-Play Alignment of Large Language Models

It is shown that RSPO with appropriate regularizers can substantially improve the length-controlled win rate on AlpacaEval-2 across a range of base models, while also achieving consistently superior performance on Arena-Hard, MT-Bench, ArmoRM, and response diversity.

Xiaohang Tang, Sangwoong Yoon, Seongho Son et al. · 6 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.