A large-scale empirical study of six zero-cost Query Performance Prediction metrics makes them a practical zero-cost replacement for magnitude thresholding in deployed RAG systems, making them a practical zero-cost replacement for magnitude thresholding in deployed RAG systems.
Abstract
Many production RAG systems implement retrieval abstention by thresholding raw similarity scores, implicitly treating score magnitude as a confidence signal. We demonstrate that this practice degrades systematically as queries require reasoning beyond semantic matching. Across 11 retrieval architectures and 28 datasets, neural retrievers consistently assign high similarity scores to semantically related but constraint-violating documents, causing magnitude-based thresholds to collapse toward near-random abstention performance on logical and temporal reasoning tasks---a failure we term the Magnitude Mirage. To address this without computationally expensive alternatives, we conduct a large-scale empirical study of six zero-cost Query Performance Prediction (QPP) metrics across three cognitive tiers: semantic matching (BEIR), logical reasoning (BRIGHT), and temporal reasoning (TEMPO). Our central finding is that the key improvement comes from abandoning magnitude in favor of score-distribution signals: the gain from this shift exceeds the differences among distributional alternatives by a factor of 5-10$\times$. In particular, Score Gap ($s_1 - s_k$) and a practical adaptation of Score Magnitude and Variance (LSMV) improve abstention AUROC by up to 0.16 in settings where magnitude-based confidence provides little discriminative power. These methods require no additional inference, retraining, or latency, making them a practical zero-cost replacement for magnitude thresholding in deployed RAG systems.
STAR is presented, a structure-aware adaptive retrieval framework for RAG that treats this mismatch as a problem of diagnosing evidence sufficiency and benefits from a control signal that preserves structurally distinct insufficiency patterns rather than collapsing them into a single scalar confidence estimate.
Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirect evidence, such as a table cell whose meaning depends on its row, column, datatype, and context. We call this mismatch the retrieval readiness gap. Our analysis shows that the curr...
Lin-Hai Ma, Ethan F. Wei, Xue-Qing Peng et al.· 0 citations
A Bayesian evaluation framework is introduced that jointly models retrieval success, abstention behavior, and answer correctness, factorized according to the pipeline's information flow, and extends to incorporate LLM-as-a-judge annotations as calibrated noisy observations, enabling practitioners to combine limited hum...
Pius von Däniken, Felix Matthias Saaro, Mark Cieliebak et al.· 0 citations
Standard Retrieval-Augmented Generation (RAG) pipelines often provide no reliable inference-time signal of whether retrieval succeeded; on ambiguous or out-of-scope queries, generation may then hallucinate. Motivated by a Czech nuclear-regulator deployment where data sensitivity precludes third-party LLM APIs, we compa...
Large language models can process increasingly long prompts, yet their ability to locate and use decisive evidence may degrade as irrelevant or confusable context is added. We formulate this phenomenon, which we call context poisoning, as extreme-value interference in attention: the decisive-evidence score is upper-bou...
Meysam Ghaffari, Nina Fatehi, Bhaskar Sen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.