Adaptive Curriculum RLOO for Verifiable Arithmetic Reasoning in Language Models
This project asks whether online reinforcement learning can improve an already supervised-finetuned Countdown model, and whether dynamically choosing problem difficulty can make the reward signal more useful.
Lucianna Kelechi Onuoha
· 0 citations