Reinforcement Learning for Self-Guided Context Compression in Mathematical Reasoning
This project tackles the Countdown math reasoning problem, and aims to train a model to use a tool which summarizes its own reasoning context, thereby compacting its reasoning trace.
Ryan Wang, Jeremy Wang
· 0 citations