seek, Self-Evaluative Exploration for Knowledge Retrieval, a training-free framework that addresses this limitation through iterative corpus interaction at test time through iterative corpus interaction at test time.
QueryRoute is introduced, a benchmark that freezes the expensive artifacts needed to study this inference-time decision problem reproducibly: original queries, generated variants, ranked lists under multiple retrievers, retrieval scores, and per-query oracle labels.
Hai-Son Le, Negar Arabzadeh, Amin Bigdeli et al.· 0 citations
Deep research agents answer complex questions through iterative loops of searching, reading, and reasoning. Recent work on reasoning-intensive benchmarks such as BrowseComp-Plus shows that well-configured lexical retrieval can surface high-quality evidence, yet agents may still fail to connect documents carrying eviden...