Skip to content

Author

Himani S. Deshpande

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#reinforcement learning Open access Oct 2026

Step-Level Gradient Masking for GRPO: Selective Optimization of Reasoning Trajectories

Group Relative Policy Optimization (GRPO) has recently emerged as an effective algorithm for Reinforcement Learning from Verifiable Rewards (RLVR), allowing for improvements in logical reasoning, more specifically in mathematical reasoning in large language models (LLMs) without supervised reasoning traces. It does, ho...

Youhan Lalwani, Mahesh Patel, Himani S. Deshpande · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.