Jul 2026
Measuring Reward-Seeking via Contrastive Belief Updates
The results indicate that RL can increase reward-seeking over the course of training, producing models that may act against their developers'intentions when they believe that doing so leads to higher reward.
Axel Højmark, Jérémy Scheurer, Evgenia Nitishinskaya et al.
· arXiv.org · 2 citations