Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval
SaMer is proposed, an object-aware token merging framework that compresses image-side post-projector tokens into representative centroids while preserving the original late-interaction interface, and outperforms compression baselines and shows stronger phrase-level grounding, suggesting that efficient multi-vector retrieval depends not only on reducing token count, but on preserving the evidence future query tokens need to select.