Skip to content

Author

Wenqian Xing

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

Attention Limited Reward Learning

Pairwise human comparisons are a primary interface through which modern AI systems learn human preferences. RLHF and related alignment pipelines typically model such comparisons with Bradley--Terry log-odds, where choice probabilities are governed by latent reward differences. This paper examines what this assumption m...

Wenqian Xing · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.