Preprint
Aug 2026
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval
This study provides the first direct comparison of native multimodal embeddings against LLM-based visual ranking on Flickr30k, and observes that GPT-4.1 and Claude Sonnet 4.6 perform on par with Gemini Embedding 2.
Archana Dutta, Vyanktesh Kanungo
· 0 citations