ITER, an agent interaction-aware dense retriever trained using agent trajectory learning signals, is introduced andlations show that structured interaction history and pre-search reasoning provide complementary retrieval context, while previously visited and useful documents provide the strongest trajectory-relative supervision.
Abstract
Deep-research agents answer complex user questions through an iterative sequence of search steps, where the agent autonomously formulates sub-queries to retrieve the evidence needed at each stage. However, existing retriever training typically relies only on the sub-query and its corresponding search results at the current step as training signals, leaving the information accumulated from previous interactions largely underutilized. We introduce ITER, an agent interaction-aware dense retriever trained using agent trajectory learning signals. ITER represents each query by incorporating not only the current sub-query, but also the main question, the agent's pre-search reasoning, and preceding sub-queries, and is trained using trajectory-relative learning signals derived from the agent's interactions. Across six agent backbones from three model families, ITER consistently outperforms the existing agent-trajectory-trained dense retriever, LRAT, achieving an average relative improvement of 6.9% on InfoSeek-Eval and 15.4% on BrowseComp-Plus. At the matched 4B scale, ITER outperforms AgentIR on InfoSeek-Eval for five of six backbones while achieving a higher visit-to-search recall ratio on BrowseComp-Plus across all six backbones. Ablations further show that structured interaction history and pre-search reasoning provide complementary retrieval context, while previously visited and useful documents, used as redundancy negatives in subsequent searches, provide the strongest trajectory-relative supervision.
A set of unsupervised signals and a Retrieval-Aware Agent Controller (RAAC) are introduced, which assists the agent in selecting optimal actions at each stage of the research process, resulting in more effective reasoning trajectories that improve overall performance while reducing unnecessary iterations, and consequen...
Heydar Soudani, Elisabeth Lingg, Faegheh Hasibi et al.· 1 citation
Interactive retrieval with partial evidence constitutes a sequential information-acquisition problem: an agent must choose questions that acquire useful evidence for the next retrieval update. Existing systems train this decision by imitating an offline ordering of candidate QA pairs. However, a question's value depend...
Lyucheng Qian, John Yuehan Zhang, Ping-Yu Wang· 0 citations
WeAgent-Harness, a multimodal agentic harness that supports native text-vision interaction and runtime recovery, and WeAgent-MMSearch, an integrated system spanning data construction, agentic post-training, and multimodal rollout that outperform similarly sized open-source models and rival models with roughly ten times...
Zongkai Liu, Hui Zhang, Li-Qiang Niu et al.· 0 citations
Existing long-video agents acquire evidence through one uniform behavior, ignoring whether the required evidence is concentrated, requires broad occurrence coverage, or must discriminate competing hypotheses---which can cause failure before substantive reasoning begins. Prescribing a fine-grained solution procedure for...
Can-Can Zhang, Bao-Feng Zhang, Xiao-Tian Han et al.· 0 citations
Deep research agents answer complex questions through iterative loops of searching, reading, and reasoning. Recent work on reasoning-intensive benchmarks such as BrowseComp-Plus shows that well-configured lexical retrieval can surface high-quality evidence, yet agents may still fail to connect documents carrying eviden...
LLMs enable multi-agent systems (MAS) to tackle complex tasks, but manually designing agent roles, prompts, and communication structures requires substantial expertise and effort. This motivates learning policies that construct query-specific MAS from execution reward. Existing approaches typically train these policies...
Bei-Cheng Xu, Bo-Wen Fan, Wei Qian et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.