Skip to content

Author

Youhan Lalwani

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#reinforcement learning Open access Oct 2026

PANES: Policy-guided Asymmetric Nash Equilibrium Search

Trick-taking card games with mandatory bidding confront reinforcement learning agents with a distinctive two-phase problem: each player must commit to a numeric bid before any cards are played, and whether that bid turns out to be correct hinges on adversarial interactions unfolding over many subsequent tricks. Judgeme...

Youhan Lalwani, Mahesh Patel, Parthavi Gaikwad et al. · 0 citations
#reinforcement learning Open access Oct 2026

Step-Level Gradient Masking for GRPO: Selective Optimization of Reasoning Trajectories

Group Relative Policy Optimization (GRPO) has recently emerged as an effective algorithm for Reinforcement Learning from Verifiable Rewards (RLVR), allowing for improvements in logical reasoning, more specifically in mathematical reasoning in large language models (LLMs) without supervised reasoning traces. It does, ho...

Youhan Lalwani, Mahesh Patel, Himani S. Deshpande · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.