PivoARL is proposed, a self-feedback retry framework for experience exploitation in LLM agents that identifies the pivotal erroneous turn through structured reflection and performs local retry only from the corresponding pivotal state, thereby reusing the correct prefix and reducing redundant interactions.
IBA-Bench is introduced, a benchmark for implicit behavioral alignment constructed from longitudinal interaction histories that contain noise, implicit cues, and temporal inconsistencies, and the proposed IBA-Agent is proposed, an agent framework that reconciles conflicting priorities through broad retrieval and trajectory-level alignment.
Jiajia Song, Bobo Li, Haiwen Yi et al.· 0 citations
This work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.
Dongfang Li, Xiaodong Luo, Ruoyu Sun et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.