Jul 2026· 55th International Conference on Environmental Systems· 0 citations
TL;DR
Results demonstrate that hypothetical document generation does not improve ranking performance for short-form or long-form queries, and a Hypothetical Document Embeddings (HyDE) layer into the retrieval pipeline is proposed and evaluated.
Abstract
This paper builds directly on our prior work presented at ICES 2025, which introduced a hybrid retrieval system combining sparse lexical retrieval (BM25), dense vector embeddings, and Reciprocal Rank Fusion to streamline NASA research search and retrieval. In that work, experimental results revealed a critical limitation: short-form queries, typical of user search behavior, consistently underperformed longer, semantically-rich queries across all retrieval strategies, with sparse retrieval exhibiting the largest degradation. This performance gap was attributed to a fundamental mismatch between the brevity of user queries and the length, jargon density, and conceptual structure of technical abstracts within the corpus. To address this limitation, this work proposes and evaluates the integration of a Hypothetical Document Embeddings (HyDE) layer into the retrieval pipeline. Rather than embedding the raw user query, the system first generates a hypothetical abstract that reflects the theoretical content, terminology, and structure of a relevant technical paper's abstract, corresponding to the intent of the user’s query. This generated abstract is then used as the retrieval query for both sparse and dense search methods. By increasing semantic density and domain-specific language, the hypothetical document theoretically improves alignment with indexed abstracts in both term-frequency and embedding space. We integrate this HyDE-based approach into the existing modular hybrid retrieval architecture and evaluate its impact on retrieval effectiveness across varying query lengths and retrieval strategies. Contrary to expectations, the results demonstrate that hypothetical document generation does not improve ranking performance for short-form or long-form queries.
RATIO (Retrieval Across Typed Ideation Operations), a large-scale benchmark in which relevance is defined by three operations which are name ideation moves, provides a scalable training and evaluation framework for retrieval components that support literature-grounded ideation, opening up new research avenues on scient...
The Adaptive Multi-Stage Vector Retrieval (AMSVR) framework is proposed, prioritising weighted, drift-resistant composition over uniform fusion, and offers tailored configurations: AMSVR-Scientific (dense + tuned hybrid) peaks at NDCG@10 = 0.7570 on SciFact, while AMSVR-Full (seven stages) targets broader, noisier corp...
Samsudeen Alabi Bankole, Yakub Kayode Saheed· NLP & Big Data· 0 citations
This work introduces AnchorQE, a training-free method that separately encodes the original query and its expansion before interpolating them, and shows that AnchorQE improves retrieval effectiveness by up to 12.89% when compared to widely-used expansion-only or text-level concatenation baselines across TREC-DL, LoTTE,...
We introduce SciRet, a compute-aware empirical study of retrieval-augmented generation for scientific question answering over CORD-19. Rather than proposing a new model, we evaluate a fixed scientific RAG pipeline across three corpus scales: 1,034 chunks (1K papers), 5,160 chunks (5K papers), and 15,480 chunks (15K pap...
Kaysarul Anas Apurba, Mahade Hasan, Rofiqul Alam Shehab et al.· 0 citations
Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirect evidence, such as a table cell whose meaning depends on its row, column, datatype, and context. We call this mismatch the retrieval readiness gap. Our analysis shows that the curr...
Lin-Hai Ma, Ethan F. Wei, Xue-Qing Peng et al.· 0 citations
This work presents Cross-Encoder Query Expansion (CE-QE), which reads the per-token relevance attributions of a cross-encoder applied to top semantic search results, selects the terms the cross-encoder treats as decisive, and appends them to the BM25 query.
Adam Kahirov, Umesh Deshpande, S. Sundararaman· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.