People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Re...
Jonathan Light, C. Cui, Jeonghye Kim et al.· 0 citations
On-policy self-distillation (OPSD) provides dense teacher targets, but evaluates them only along student-sampled rollouts. When the privileged teacher favors an alternative action at a visited prefix, OPSD can provide a target for the branch decision but cannot supervise the successor contexts induced by that action un...
Taeckyung Lee, Rinat Amankos, Jeonghye Kim et al.· 0 citations
ProgramDistill is introduced, a benchmark evaluating coding agents on features discovered through interaction with fully functional reference applications, and provides a scalable benchmark with controlled difficulty for evaluating and diagnosing coding agents, and a natural basis for future curriculum-based training.
Jeonghye Kim, Minseon Kim, Young Jin Kim et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.