This paper presents a controlled analysis of whether spread-out regularization, a batch-level loss that penalizes high pairwise similarity within a batch, can improve representation-space utilization in Matryoshka Representation Learning settings and shows that spread-out loss is a retrieval-biased regularizer.
Text embeddings from retrieval-tuned (dual-encoder) models are increasingly used as context features in contextual bandits for recommendation, on the assumption that an embedding space optimized for inner-product similarity will speed up a linear exploration policy. This study tests that assumption with a controlled, s...
HN-CLIP is introduced, which uses the text encoder's own text-text geometry to construct per-negative adaptive similarity margins, and improves all six tested fine-tuning frameworks on the in-domain benchmarks and reaches the strongest full-data baseline with only 20% of the training data.
Hao-Yue Liu, Ye-Heng Chen, Zhi-Chao Wang et al.· 0 citations
This work provides the first explicit family of query and document sets, together with their relevance matrices, for which single-vector embeddings that rank all relevant documents above irrelevant ones require exponential size, whereas polynomial-size multi-vector embeddings suffice.
Mihir Agarwal, Viraj Agrawal, Sabyasachi Basu et al.· 1 citation
DESS is introduced, a lightweight uncertainty layer that augments an existing embedding model with a predicted mean vector and an independent per-dimension spread vector that provides a modular, geometry-aware uncertainty layer for embedding-space models, provided its spread is calibrated to local embedding geometry.
Morten Grundetjern, J. Voigt, Per-Arne Andersen et al.· KI - Künstliche Intelligenz· 0 citations
This work introduces Giga-Embeddings, a family of text embedding models designed to combine strong retrieval quality with efficient serving, and trains the compact model using a dimension-agnostic objective that aligns teacher and student similarity distributions.
Egor Kolodin, Egor Krasnoperov, Evgeniy Kosarev et al.· 0 citations
Multi-vector visual document retrieval (VDR) models such as ColPali and ColNomic achieve strong accuracy by representing each document with hundreds to thousands of patch-level embeddings, at substantial storage and latency cost. Existing compression methods either prune unimportant patches or merge similar ones into c...
Jian-Xin You, Kun Ni· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.