DCE+SRCL outperforms OPSD across multiple model scales and four competition-level mathematics benchmarks, and its comprehensive evaluations show that DCE+SRCL outperforms OPSD across multiple model scales and four competition-level mathematics benchmarks.
Shang-Jian Yin, Ze-Hao Zhao, Kavosh Asadi et al.· 0 citations
Off-Context GRPO (OC-GRPO), a minimally modified variant of GRPO that uses guided rollouts but applies an importance-corrected objective to steer the update back toward the original unguided objective, avoiding the mismatch that destabilizes uncorrected guided training.
Priyank Agrawal, Ankur Samanta, S. Ghasemlou et al.· arXiv.org· 1 citation
Large Language Models (LLMs) have emerged as powerful assets for recommender systems. However, deploying them as generative recommenders or zero-shot rankers at web-scale remains bottlenecked by prohibitive computational overhead and grounding challenges. In this paper, we revitalize the classic, highly efficient two-t...