Skip to content

Author

Aizada Nurdinova

We have 1 of 1 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

SFT Augmentation and Replay-Based RL for Countdown Reasoning

This work studies how supervised fine-tuning data quality and reinforcement learning (RL) optimization design affect performance on Countdown arithmetic reasoning under a reinforcement learning with verifiable rewards (RLVR) setup, and investigates whether replay-based reinforcement learning can improve sample efficiency during RL fine-tuning.

Adhi Daiv, Aizada Nurdinova, Ellie Sampson · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.