Beyond Pairwise Feedback: Listwise Vision-Language Supervision for Preference-Based Reward Learning
It is shown that Plackett-Luce (PL) reward models can train robotic policies from VLM-generated rankings as effectively as pairwise Bradley-Terry, $K$-wise Bradley-Terry, and RL-VLM-F baselines and demonstrate that listwise VLM preference supervision is a competitive and flexible approach to reward learning for reinfor...