Skip to content

One Size Does Not Fit All! Dynamic Retriever and Generator Selection for RAG

Sep 2026 · 0 citations · 70 references
Computer Science

TL;DR

This work systematically analyze how retriever and generator complexity interacts across factoid and multi-hop question answering (QA), including bridge and composition reasoning tasks, and introduces DRAG, a query-adaptive framework for selecting retriever-generator configurations.

Abstract

Retrieval-Augmented Generation (RAG) systems typically employ fixed retriever and generator configurations across queries, despite substantial differences in query complexity and information needs, leading to inefficient allocation of computational resources. While retrieval and generation adaptivity have been studied independently, their joint effect on end-to-end RAG performance remains underexplored. We systematically analyze how retriever and generator complexity interacts across factoid and multi-hop question answering (QA), including bridge and composition reasoning tasks. Our analysis shows that stronger retrieval generally yields larger gains than increased generation effort, but both exhibit diminishing and non-monotonic returns, indicating that higher-complexity configurations are not uniformly better across queries. Motivated by these findings, we introduce DRAG, a query-adaptive framework for selecting retriever-generator configurations. We first propose DRAG$_\text{QPP}$, a training-free routing approach that uses Query Performance Prediction (QPP) signals to guide retriever selection and perplexity-based measures over retrieved context to guide generator selection. We further introduce DRAG$_\text{SFT}$, a supervised routing approach that fine-tunes an LLM to jointly predict retriever-generator configurations. Across three LLM families and four QA benchmarks, \qpprag~achieves performance comparable to strong static RAG baselines while substantially reducing inference latency, whereas DRAG$_\text{SFT}$ consistently improves effectiveness over static and training-free adaptive baselines. Overall, DRAG demonstrates that jointly adapting retrieval and generation achieves a more favorable effectiveness-efficiency trade-off than static RAG pipelines.

View source

Similar papers

Preprint Sep 2026

Route Me If You Can: A Benchmark for Query Reformulation Selection

QueryRoute is introduced, a benchmark that freezes the expensive artifacts needed to study this inference-time decision problem reproducibly: original queries, generated variants, ranked lists under multiple retrievers, retrieval scores, and per-query oracle labels.

Hai-Son Le, Negar Arabzadeh, Amin Bigdeli et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LiteRAG: Cost-Efficient Graph-Based Retrieval-Augmented Generation

Graph-based retrieval can improve multi-hop question answering, but existing approaches often incur high query-time costs and produce diffuse, oversized contexts that reduce generation efficiency. We present LiteRAG, a graph-based retrieval method that replaces expensive retrieval-time LLM control with query-conditione...

Daniel Alejandro Coll Tejeda, Pedro García López, Daniel Barcelona-Pons · 0 citations

STAR: Structure-Aware Adaptive Retrieval for RAG

STAR is presented, a structure-aware adaptive retrieval framework for RAG that treats this mismatch as a problem of diagnosing evidence sufficiency and benefits from a control signal that preserves structurally distinct insufficiency patterns rather than collapsing them into a single scalar confidence estimate.

Yeowon Jeon, Chong-kwon Kim, Y. Choi · 0 citations
Review Open access Sep 2026

Energy-Efficient Retrieval-Augmented Generation: A Systematic Review of Retrieval, Reranking, Caching, and Context-Management Strategies

Retrieval-Augmented Generation (RAG) has become a standard paradigm for grounding large language models (LLMs) in external knowledge, but typical deployments introduce substantial energy, latency, and cost overhead due to expensive retrieval and context-processing pipelines. Recent work in "Green AI" and sustainable ma...

Anupam Dhakal, Prashant Pokharel, S. Adhikari · 0 citations
#artificial intelligence Preprint Sep 2026

ORDER: Task-Conditioned Routing for Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) pipelines typically rely on a fixed indexing and retrieval configuration determined at preprocessing time. This one-size-fits-all design is ill-suited to domain-expert settings, where heterogeneous queries require different chunking granularities, metadata constraints, and source-se...

Aurélien Pellet, Julien Perez, Marie Puren · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond the Query: Do Retrieval Signals Improve Adaptive Multimodal RAG Routing?

Adaptive RAG often uses retrieval-time signals to decide whether another retrieval, reranking, or multimodal step should run. We ask whether these signals add routing value once the query itself is already known. Across document, audio, and video RAG, we compare matched query-only and query+retrieval routers while hold...

Qiao-Mu Li, Qiu-Yuan Zhang, Nong Ming · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.