#machine learning
Jul 2026
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information
Off-Context GRPO (OC-GRPO), a minimally modified variant of GRPO that uses guided rollouts but applies an importance-corrected objective to steer the update back toward the original unguided objective, avoiding the mismatch that destabilizes uncorrected guided training.
Priyank Agrawal, Ankur Samanta, S. Ghasemlou et al.
· arXiv.org · 1 citation