Reflection on Visual Evidence (ReVuE), an on-policy distillation method for visual agents that translates trajectory-level evidence diagnosis into targeted token-level supervision, guiding students to improve their visual evidence acquisition and reasoning.
Shaohang Wei, Feifan Song, Guangyue Peng et al.· 0 citations
It is shown that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives too rare to sample and reinforce, and effective rewardable support is defined as successful trajectories reachable within a fixed rollout budget.
Shaohang Wei, Z.Y. Su, Feifan Song et al.· 0 citations
This work proposes Experience-Augmented Policy Optimization (EAPO), which leverages a prior RL-optimized policy as an action-level experience prior and selectively injects experience at critical decision points during rollout to ensure stable and unbiased learning from experience-augmented rollouts.