Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or leave the benefit of ex...
Zi-Han Lin, Xiao-Han Wang, Jie Cao et al.· 0 citations
Reinforcement learning (RL) has become a central paradigm for large language model (LLM) post-training, but optimization toward new objectives can degrade capabilities already present in the base model. KL regularization is widely used to mitigate such forgetting by constraining policy drift toward a reference model. H...
Li Wang, Xiaodong Lu, Xiaohan Wang et al.· 0 citations
HiDiffTIR is proposed, a Hierarchical Difficulty-aware policy optimization framework for multi-turn TIR that consistently improves multi-turn TIR performance and tool invocation accuracy over strong RL baselines, highlighting the necessity of difficulty-aware credit assignment for effective policy optimization in tool-...
Yu-Can Guo, Xiaohan Wang, Miao Su et al.· 0 citations
BASM is proposed, which augments each skill with explicit boundary fields, which transforms each retrieved skill from an unconditional action template into state-conditioned guidance: the agent applies the skill when its conditions hold, suppresses inapplicable tool calls when they do not, and issues targeted repairs w...
Zi-Han Lin, Zhenyu Chen, Jiawen Wei et al.· 0 citations
Inspired by the human brain, which balances plasticity and stability through complementary episodic storage and gradual consolidation, UniMem is proposed, a self-routing framework for autonomous memory management that consistently outperforms baselines while maintaining execution fidelity.