Skip to content

Author

Jiajun Chai

We have 3 of 33 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning

Tool-Integrated Reasoning (TIR) is a fundamental capability for LLM agents to solve complex tasks by interacting with external tools iteratively. Reinforcement Learning (RL) has become the dominant paradigm for enabling this capability. However, existing approaches typically assign uniform trajectory-level advantages and treat all correct tool calls equally, ignoring the varying difficulty and learning value across trajectories and reasoning steps. This can lead to imprecise learning signals that do not adequately distinguish between trivial and challenging tool-use patterns. To address this limitation, we propose HiDiffTIR, a Hierarchical Difficulty-aware policy optimization framework for multi-turn TIR. HiDiffTIR performs difficulty-aware credit assignment at both trajectory and turn levels, enabling the policy to focus on more informative trajectories and harder reasoning steps. Notably, this fine-grained optimization is achieved without additional supervision, relying solely on group-level statistics derived from standard RL rollouts. Extensive experiments on three tool-using benchmarks demonstrate that HiDiffTIR consistently improves multi-turn TIR performance and tool invocation accuracy over strong RL baselines, highlighting the necessity of difficulty-aware credit assignment for effective policy optimization in tool-integrated LLM agents.

Yucan Guo, Xiaohan Wang, Miao Su et al. · 0 citations
Preprint Aug 2026

When Not to Imitate: Boundary-Aware Skill Memory for Reliable Tool-Use LLM Agents

BASM is proposed, which augments each skill with explicit boundary fields, which transforms each retrieved skill from an unconditional action template into state-conditioned guidance: the agent applies the skill when its conditions hold, suppresses inapplicable tool calls when they do not, and issues targeted repairs when execution fails.

Zi-Han Lin, Zhenyu Chen, Jiawen Wei et al. · 0 citations
Jul 2026

UniMem: Complementary Episodic-to-Parametric Memory for Boundary-Agnostic Task Streams

Inspired by the human brain, which balances plasticity and stability through complementary episodic storage and gradual consolidation, UniMem is proposed, a self-routing framework for autonomous memory management that consistently outperforms baselines while maintaining execution fidelity.

Si-Yu Xia, Chen-Heng Zhang, Yan-Ting Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.