Skip to content
Preprint

Route Me If You Can: A Benchmark for Query Reformulation Selection

Sep 2026 · 0 citations · 45 references
Computer Science

TL;DR

QueryRoute is introduced, a benchmark that freezes the expensive artifacts needed to study this inference-time decision problem reproducibly: original queries, generated variants, ranked lists under multiple retrievers, retrieval scores, and per-query oracle labels.

Abstract

LLM-based query reformulation can improve retrieval, but no single reformulation strategy is consistently optimal across queries, domains, retrievers, or model backbones. This creates an inference-time decision problem: ``Given an original query and a pool of candidate reformulations, which one should be issued to the retriever?''. Existing studies are hard to compare because they use different reformulator pools, retrievers, relevance signals, training labels, and evaluation metrics. We introduce QueryRoute, a benchmark that freezes the expensive artifacts needed to study this decision reproducibly: original queries, generated variants, ranked lists under multiple retrievers, retrieval scores, and per-query oracle labels. The benchmark contains 3,757 queries, 11 candidate systems, five reformulator backbones, and three retrievers across TREC DL, BEIR, and BRIGHT, yielding 619,905 retrieval outcomes. We benchmark supervised classification, routing, QPP, and LLM-as-judge selectors. Results show substantial oracle headroom over fixed reformulators, but current selectors recover only part of it; selector rankings change across retrievers, and similar mean effectiveness can hide different query-level behavior. The released artifacts and evaluation harness allow future selectors to be compared without regenerating variants, rerunning retrieval, or rebuilding judge pipelines. Code and data are available at https://github.com/haisonle001/QueryRoute

View source

Similar papers

#artificial intelligence Preprint Sep 2026

One Size Does Not Fit All! Dynamic Retriever and Generator Selection for RAG

This work systematically analyze how retriever and generator complexity interacts across factoid and multi-hop question answering (QA), including bridge and composition reasoning tasks, and introduces DRAG, a query-adaptive framework for selecting retriever-generator configurations.

Neeraj Anand, Payel Santra, Partha Basuchowdhuri et al. · 0 citations
Preprint Aug 2026

EXPLAIN Yourself! Finding Query Planner Stalls Across DBMSes

It is found that although the queries triggering slow planning are largely DBMS-specific, recurring pathologies involving correlated subqueries, CTE expansion, repeated subquery expressions, disjunctive joins, and constant folding affect multiple systems.

Geoffrey X. Yu, Ryan Marcus, Tim Kraska · 0 citations
Preprint Sep 2026

Pre-retrieval Query Clustering for Adaptive Top-k Document Retrieval in RAG Systems

RAG systems commonly retrieve a fixed number of documents (top-k) to ground generation, but this static approach is brittle: simple queries suffer over-retrieval (adding noise and cost) while complex queries are under-retrieved, causing recall failures that cascade into incorrect answers. Motivated by the question of h...

Ye Xia, Emre Yamangil, Hai-Xun Wang · 0 citations
#natural language process... Preprint Sep 2026

Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems

Evaluating first-stage retrievers in large-scale production RAG requires a benchmark that pairs a large-scale corpus with a large set of agent-reformulated search queries based on real user queries and their conversation threads, and that labels many relevant documents per query. No existing public benchmark evaluates...

Maximilian Schall, Sedigheh Eslami, Markus Krimmel et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond the Query: Do Retrieval Signals Improve Adaptive Multimodal RAG Routing?

Adaptive RAG often uses retrieval-time signals to decide whether another retrieval, reranking, or multimodal step should run. We ask whether these signals add routing value once the query itself is already known. Across document, audio, and video RAG, we compare matched query-only and query+retrieval routers while hold...

Qiao-Mu Li, Qiu-Yuan Zhang, Nong Ming · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.