#artificial intelligence
May 2026
RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking
RankQ, an offline-to-online Q-learning objective that augments temporal-difference learning with a self-supervised multi-term ranking loss to enforce structured action ordering is proposed, which shapes the Q-function such that action gradients are directed toward higher-quality behaviors.
Andrew Choi, Wei Xu
· arXiv.org · 2 citations