Skip to content

Author

Md. Rezwanul Haque

We have 1 of 9 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

Towards Robust Reinforcement Learning for Small-Scale Language Model Agents

A capacity-headroom hypothesis is proposed, which states that PPO performance at the SLM scale depends on both a fluent supervised model and a discriminative reward signal, rather than on the number of model parameters.

Md. Rezwanul Haque, Md. Milon Islam, Fakhri Karray · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.