Jul 2026
Rewarding Better Thinking for LLM Preference Alignment
Think Checklist Reward (TCR), a process-oriented reward for RL-based preference alignment that converts preference pairs into sample-specific thinking checklists and uses them to evaluate whether the generated reasoning trace addresses the preference-implied considerations.
Xu-Bo Liu, Wenya Guo, Ruxue Yan et al.
· arXiv.org · 0 citations