Skip to content

Author

Ankur Samanta

We have 2 of 14 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

Off-Context GRPO (OC-GRPO), a minimally modified variant of GRPO that uses guided rollouts but applies an importance-corrected objective to steer the update back toward the original unguided objective, avoiding the mismatch that destabilizes uncorrected guided training.

Priyank Agrawal, Ankur Samanta, S. Ghasemlou et al. · 1 citation
Jul 2026

Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias

This work develops a black-box method for learning a single context-independent logit-bias vector, added at every decoding step, without modifying model weights or requiring gradients, and suggests that learned logit bias is a lightweight mechanism for adapting language models under minimal access requirements.

Ofek Cohen, Lior Shani, Aviv Rosenberg et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.