At test time, large language models (LLMs) can encode historical information in activation memory (i.e., KV caches) and parametric memory (i.e., updated parameters). While activation memory is generally considered effective for factual recall and parametric memory for learning new tasks, their interplay remains unclear...
Miao-He Niu, Run-Song Zhao, Xin-Yu Liu et al.· 0 citations
Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models. However, despite their strong capabilities in response ranking, generative reward models have not realized their potential in reinforcement learning (RL). Our analysis reveals that this limitation ari...
Chenglong Wang, Ziming Zhu, Yifu Huo et al.· 1 citation
Environmental Feedback-based Credit Assignment (EFCA), a multi-timescale credit assignment approach for long-horizon agentic RL that complements the long-term outcome signal with two environment-grounded process signals: a short-term feedback signal that captures the immediate effect of the current action and a medium-...
Yifu Huo, Shunjie Xing, Chenglong Wang et al.· 1 citation
Flow Continuous Trajectory Supervision (FlowCTS), which matches subsequent student and reference trajectories initialized from the same student-visited state to derive a temporally weighted velocity-matching upper bound and discretize it into practical objectives parameterized by the number of supervision steps.
Kaiyang Ye, Yuan Ge, Junxia Zhang et al.· arXiv.org· 0 citations
The method Syfer is introduced, a synthesizer-folding framework for multilingual multi-hop question answering that defers translation rather than applying it by default and attains competitive accuracy while striking a favourable balance between performance and computational cost.
Yilin Wang, Yuchun Fan, Weidong Bao et al.· 0 citations
This work proposes Dynamic Decomposition and Filtering for Multi-Hop Reasoning-Augmented Generation (D2F-ReAG), a novel paradigm that adaptively controls reasoning depth by judging the reliability of the root-level reasoning.
Jiaoyang Li, Junhao Ruan, Sheng-Wei Tang et al.· 0 citations
ToFu is presented, an agentic harness for researchers that reads your codebase, edits files, runs commands, and integrates with your development tools and provides a white-box agentic harness that allows researchers to inspect, modify, and evaluate its orchestration logic, tool-use behavior, and harness design.
Junhao Ruan, Yuan Ge, Bei Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.