Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control
It is shown that selective control under partial auditing reduces accepted errors while increasing correctness, and experiments with log linear and neural contextual bandits and with a language model support the analysis and show that selective control under partial auditing reduces accepted errors while increasing cor...