Skip to content
Book Open access

PrepRet: Automated Data Preparation Pipeline Selection for Neural Retrieval

Jul 2026 · Annual International ACM SIGIR Conference on Research and Development in Information Retrieval · pp. 3642-3646 · 0 citations · 15 references
Computer Science

TL;DR

PreRet, a framework that jointly optimizes preprocessing pipeline selection and neural retrieval through differentiable optimization through differentiable optimization, is proposed, enabling end-to-end gradient-based learning over a search space of 700 configurations spanning text cleaning, pre-tokenization, chunking, and augmentation.

Abstract

Neural retrieval models typically rely on fixed, hand-crafted preprocessing pipelines designed independently of the retrieval task, leading to suboptimal performance that varies across datasets and architectures. We propose PrepRet, a framework that jointly optimizes preprocessing pipeline selection and neural retrieval through differentiable optimization. We formulate preprocessing selection as a differentiable discrete choice problem using Gumbel-Softmax relaxation, enabling end-to-end gradient-based learning over a search space of 700 configurations spanning text cleaning, pre-tokenization, chunking, and augmentation. A hierarchical selection mechanism captures inter-stage dependencies between preprocessing operations. On MS MARCO, PrepRet achieves 0.533 nDCG@10, improving over Grid Search by 7.7% and Contriever by 11.7%, while requiring only 1.1 GPU-hours. Zero-shot evaluation on eight BEIR datasets confirms robust cross-domain generalization, with particularly strong gains on scientific and entity-rich domains.

Read PDF

Similar papers

Open access Jul 2026

ADAPTIVE MULTI-STAGE VECTOR RETRIEVAL FOR RETRIEVAL-AUGMENTED GENERATION

The Adaptive Multi-Stage Vector Retrieval (AMSVR) framework is proposed, prioritising weighted, drift-resistant composition over uniform fusion, and offers tailored configurations: AMSVR-Scientific (dense + tuned hybrid) peaks at NDCG@10 = 0.7570 on SciFact, while AMSVR-Full (seven stages) targets broader, noisier corpora where Recall@100 matters most.

Samsudeen Alabi Bankole, Yakub Kayode Saheed · 0 citations
Conference Jul 2026

DART: dynamic adapter refinement at test-time for multimodal document retrieval

Empirical evaluations across a diverse suite of multimodal document retrieval benchmarks reveal that DART achieves consistent and significant gains in ranking precision, and this dynamic refinement process introduces minimal computational latency, offering a highly efficient, plug-and-play solution for adaptive document retrieval.

Jing Zhang, Yaowei Wang, Chongyu Wang et al. · 0 citations
Preprint Sep 2026

REDSI: Addressing the Reproducibility and Evaluation Consistency of Differentiable Search Indexing for Document Retrieval

The differentiable search index (DSI) framework (Tay et al., 2022) has become the de facto baseline for generative retrieval. However, DSI is hard to reproduce: no public implementation covers all three original document identifier types (atomic, naive, semantic), reported results vary widely, and the ubiquitous NQ320K dataset is built from Natural Questions through diverse and underspecified preprocessing. We introduce ReDSI, the first open-source DSI implementation supporting all three identifier types, together with a parameterizable and well-documented NQ320K construction pipeline. Experimentally, we achieve results that are competitive with or stronger than previous DSI baselines. Moreover, we conduct extensive experiments under model downscaling, covering retrieval effectiveness, parameter efficiency, training methods and decoding strategies, opening novel directions for future research.

Unknown authors · 0 citations
Preprint Aug 2026

Learning Sample-wise Rank-aware Interpolation Weights for Composed Visual Data Retrieval

This work revisits the efficacy of simple linear interpolation within an embedding space, and introduces SRAIN, the first framework that dynamically predicts query-specific interpolation weights, and achieves the best in composed video retrieval and matches the current state of the art in composed image retrieval.

Boseung Jeong, T. Park, Donghyeon Kwon et al. · 1 citation
Book Open access Jul 2026

Corpus-Centric Learning for Zero-Shot Table Retrieval

GeCo-TR is proposed, a zero-shot table retrieval framework that eliminates the need for supervised QA data by shifting from direct query-to-table learning to modeling the intrinsic structural semantics of the table corpus, resulting in high-precision and high-recall retrieval for implicit queries in a zero-shot setting.

Zhou He, Zhifei Pang, Xiu Tang et al. · 0 citations
Jul 2026

SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models

It is observed that attention scores from both vision and text tokens peak at modality separator tokens, suggesting that these separators bridge the two modalities and proposes SepPrune, an efficient, training-free, plug-and-play pruning method that uses the separator token as a unified query to rank and select informative vision tokens.

Yucheng Wang, Qihui Zhu, Yang Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.