Long-horizon agents improve solutions through sustained interaction, execution, and task feedback. Scaling studies relate performance to resources and capabilities, yet how existing capabilities shape returns to extended interaction remains less understood. To address this gap, we analyze AutoLab and EdgeBench, two lon...
Hao-Yu Zheng, Zheng-Yu Chen, Huai-Sheng Zhu et al.· 0 citations
A time-truncation harness is proposed that enforces a temporal cut-off at every turn, enabling TIR-style sampling from historical events, enabling TIR-style sampling from historical events, reducing temporal leakage and reliance of rejection sampling or unsolved queries, increasing the sampling efficiency.
A systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks is presented.
Yi-Wei Li, Wanli Yang, He-Xiang Tan et al.· 1 citation
An extensive evaluation of automatic harness evolution for LLM agents is conducted, comparing harness evolution with simple test-time scaling and discovery baselines under comparable feedback and inference budgets, and also evaluating evolved harnesses on held-out tasks to assess whether the discovered improvements gen...
Yike Wang, Huaisheng Zhu, Zhengyu Hu et al.· arXiv.org· 18 citations
Student-Centric Answer Sampling (SCAS) is proposed, a framework that selects from verified teacher-generated answers according to their estimated student-centric learning cost and is derived by a token-wise gradient decomposition and used to guide answer selection during training.
Zhengyu Hu, Zheyuan Xiao, Linxin Song et al.· 0 citations
HarnessCompass is proposed, a novel automatic harness evolution framework built around constrained evolution, proactive feedback, and component-wise optimization that improves Pass@1 from 54\% to 66\% in only 5 evolution iterations, outperforming AHE in both effectiveness and evolution efficiency.
Luan Zhang, Ruo-Chen Zhou, Dan-Dan Song et al.· 10 citations· ⚡1
Co-Harness is introduced, a framework that jointly optimizes the agent harness and model parameters during post-training and suggests that joint harness and model optimization is an effective way to improve agents beyond fixed-harness post-training.
Zhengyu Chen, Teng Xiao, Huaisheng Zhu et al.· arXiv.org· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.