Long-horizon agents improve solutions through sustained interaction, execution, and task feedback. Scaling studies relate performance to resources and capabilities, yet how existing capabilities shape returns to extended interaction remains less understood. To address this gap, we analyze AutoLab and EdgeBench, two lon...
Hao-Yu Zheng, Zheng-Yu Chen, Huai-Sheng Zhu et al.· 0 citations
A time-truncation harness is proposed that enforces a temporal cut-off at every turn, enabling TIR-style sampling from historical events, enabling TIR-style sampling from historical events, reducing temporal leakage and reliance of rejection sampling or unsolved queries, increasing the sampling efficiency.
A systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks is presented.
Yi-Wei Li, Wanli Yang, He-Xiang Tan et al.· 1 citation
Co-Harness is introduced, a framework that jointly optimizes the agent harness and model parameters during post-training and suggests that joint harness and model optimization is an effective way to improve agents beyond fixed-harness post-training.
Zhengyu Chen, Teng Xiao, Huaisheng Zhu et al.· arXiv.org· 4 citations
ToFu is presented, an agentic harness for researchers that reads your codebase, edits files, runs commands, and integrates with your development tools and provides a white-box agentic harness that allows researchers to inspect, modify, and evaluate its orchestration logic, tool-use behavior, and harness design.
Junhao Ruan, Yuan Ge, Bei Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.