Skip to content

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability

Jul 2026 · arXiv.org · Vol abs/2607.11432 · 0 citations · 65 references
Computer Science

TL;DR

This work generalizes preference-based RL by formalizing a novel setting in which the expert can also label trajectory pairs as incomparable, i.e., when neither trajectory dominates the other.

Abstract

In this work, we study the reinforcement learning (RL) problem from pairwise trajectory comparisons provided by a human expert. We generalize preference-based RL by formalizing a novel setting in which the expert can also label trajectory pairs as incomparable, i.e., when neither trajectory dominates the other. We introduce the learning problem and the desiderata that its solution should satisfy. Then, we propose a novel Bradley-Terry-inspired rationality model that effectively captures incomparabilities and infers a multi-dimensional reward function, and we study its properties. We provide a sample complexity analysis for learning the model parameters when a dataset is available. Finally, we evaluate our model's ability to reconstruct a reward function that aligns with the expert's comparisons in simulated environments and to recover the Pareto frontier of policies, along with a robustness analysis across varying levels of expert rationality.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Inference-Time Nash Alignment

This work forms the problem as obtaining a Nash equilibrium of a two-player zero-sum game between policies, and proposes two algorithms: Best-of-Nash (BoN) and Nash Mirror Descent (NMD), which are proved to achieve a duality gap that matches the problem lower bound.

Hadi Hosseini, Debmalya Mandal, Duo-Han Zhang · 0 citations
Jul 2026

Reinforcement Learning: From Algorithms To Foundation Models

This thesis develops diffusion-based world models, investigates RL for efficient video generation, explores generative models as policy classes, and studies interactive video world models in which actions shape future observations, and addresses long-horizon modeling through architectures with memory.

Zihan Ding · 0 citations
Preprint Aug 2026

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling

This paper introduces a novel reward-model-free, critic-free, and gradient-based PbRL algorithm compatible with segment preferences named Segment Pairwise Proximal Policy Optimization (SP3O), and provides a theoretical basis for the algorithm and analyze the tradeoff in choosing the segment length.

Evan Assmus, Qi-Ning Zhang, Lei Ying · 0 citations
Preprint Aug 2026

Evolution of cooperation with Q-learning: how much information do we need?

Mechanistic analyses show that a moderate neighborhood size enables individuals to strike an optimal balance between information sufficiency and decision-making tractability, which allows them to detect reciprocal opportunities while avoiding the deterioration of decision quality due to information overload.

Yi-Hsin Ku, Xin Ou, Ji-Qiang Zhang et al. · 0 citations

Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning

This work proposes a game-theoretic framework that gives this reward-retention trade-off an explicit statistical interpretation, and provides a principled method for learning this equilibrium coefficient via reduction to the KL-regularized RL objective, thus allowing for flexible integration into standard fine-tuning p...

Keegan Harris, Brian Lee, Ian Waudby-Smith et al. · 0 citations
Jul 2026

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

This work introduces OVI, an interactive on-policy IL algorithm that is statistically efficient whenever the learner can represent the expert's value function and computationally efficient given access to a linear maximization oracle, and introduces a negative result showing that interaction is necessary.

Luca Viano, Antoine Moulin, Audrey Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.