A time-truncation harness is proposed that enforces a temporal cut-off at every turn, enabling TIR-style sampling from historical events, enabling TIR-style sampling from historical events, reducing temporal leakage and reliance of rejection sampling or unsolved queries, increasing the sampling efficiency.
Abstract
Future event prediction carries broad social impact yet remains challenging. SOTA approaches augment LLMs with external agent frameworks whose predictive capability vanishes once the harness is removed. While recent Tool-Integrated Reasoning (TIR) internalizes deep search for multi-hop retrieval of facts, forecasting further demands temporal search and reasoning over historical trends and dynamic shifts. The key obstacle is data: historical queries induce temporal leakage that degrades forecasting into retrieval. Prior works either freeze information gathering with static observations, or rely on rejection sampling or unresolved fresh queries that discard vast amounts of data, degrading synthesis efficiency. We propose a time-truncation harness that enforces a temporal cut-off at every turn, enabling TIR-style sampling from historical events, reducing temporal leakage and reliance of rejection sampling or unsolved queries, increasing the sampling efficiency. We further build a large-scale corpus and a process-based metric and show that our harness naturally induces a broader temporal breadth of search and raises the proportion of high-quality data, further increasing the efficiency and reducing the reliance on complex rubrics. Distillation experiments show that students trained on harness-intervened data achieve the best performance, demonstrating harness-assisted model evolving that turns higher quality temporal search and reasoning data into a parametric advancement of the students.
Deep search agents tackle complex knowledge tasks through iterative retrieval, multi-hop reasoning, and evidence synthesis across multiple sources. Existing approaches typically assume relatively stable retrieval systems and operate over short-horizon tool interaction. However, when retrieval is sensitive to query form...
Experiments establish ReTree as an effective self-correcting memory abstraction for long-horizon search, and show that ReTree consistently outperforms Full-Trajectory ReAct in question-answering and search benchmarks.
Aijun Yang, Qianxue Guo, Ziyi Huang et al.· 0 citations
This work proposes a synthetic, simulation-driven framework for studying knowledge insertion in LLMs, and introduces {\sc ParallelEvents}, a benchmark of fictional yet realistic future worlds that generates coherent event trajectories for controlled evaluation, avoiding contamination while preserving consistency.
Jonathan Zheng, Zi-Rui Shao, Alan Ritter et al.· 0 citations
Results characterize agent reasoning as dynamic working state rather than permanent interaction history, and suggest that historical reasoning becomes more replaceable once task relevant derived state has been reliably externalized into code, files, tool outputs, or environmental feedback.
Ming-Xuan Wang, Fei Luo, Bo Wang et al.· 0 citations
ThinkRetrieve is proposed, a test-time scaling framework that augments the reasoning traces of LRMs with dynamically retrieved solved examples at each reasoning step, providing the model with guidance on how to reason rather than merely what facts are relevant.