Benchmarking Pre-Trained Vision-Language Models for Bidirectional Image-Text Retrieval on MS-COCO: BERT+ResNet, CLIP, and BLIP
Bidirectional image-text retrieval evaluates whether a model can align visual and textual representations for both text-to-image and image-to-text search. This paper presents a controlled benchmark on the Microsoft Common Objects in Context (MS-COCO) Karpathy split, using the same 5,000-image test gallery, 25,000 capti...