It is argued that OPSD is a sensitive algorithm rather than a generally reliable reasoning-improvement post-training method because it produces ineffective length growth, stable degradation, or behavioral collapse.
Yang Li, Gong-Le Xue, Yu-Heng Yuan et al.· 0 citations
Modern LLM post-training composes supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and on-policy distillation (OPD) into multi-stage pipelines, yet these stages are typically designed and evaluated in isolation. We show that this composition is consequential: a stage that improves th...
Emre Can Acikgoz, Yang Li, Z. Liu et al.· 0 citations
Autoregressive Thought Flow is introduced, which models the next continuous thought as a multimodal distribution, and suggests that continuous reasoning is more effective when multiple possible next thoughts remain available rather than being collapsed into a single prediction.
Yang Li, Yi Wang, Shinda Huang et al.· 0 citations
Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively.
Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are added. We introduce EvoHarnessBench, a benchmark for evaluating agents under controlled harness evo...
Zi-Xuan Ke, Vaidehi Patil, Hai-Zhou Shi et al.· 1 citation
Experiments spanning mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks show that RISE outperforms RLVR-only training and on-policy self-distillation across all settings.
Yang Li, Semih Yavuz, Shafiq Joty· 3 citations· ⚡1
The results suggest that dense credit assignment through distillation can be effective when its likelihood-based scores are empirically validated as meaningful proxies for outcome-relevant credit, when this alignment does not hold.
Xuan-Phi Nguyen, Z. Liu, Yang Li et al.· 3 citations
Procedural Memory Distillation is proposed, which converts crossepisode signals into reusable procedural memory and distills it into the policy's weights during training, yielding a memory-free model at inference.
Ye Liu, Srijan Bansal, Bo Pang et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.