Auxiliary text-guided image restoration for image-text matching
Abstract
Image-Text Matching (ITM) aims to establish deep semantic associations between visual content and textual descriptions. Existing methods usually have discrimination issues because of fine-grained semantic deviations, so it's hard to capture the complex correspondences between cross-modal entries. Only relying on alignment loss a lot results in a lot of topranked false positives when dealing with such subtle semantic discrepancies. To resolve this issue, we introduce a new ITM framework that improves the model's discriminative performance by focusing on localized core attributes. Specifically, introducing a Semantic Augmentation Strategy (SAS) with a Text-guided Image Restoration (TIR) task, while forcing the model to extract color semantics from grayscale images, our approach effectively creates deep semantic coupling across modalities. Optimized via a multi-loss collaborative framework, AT-JLIM significantly expands the inter-class distance between positive and negative samples in the feature space. Extensive experiments on benchmark datasets, like MSCOC and Flickr30K, show that the proposed method does a lot better in fine-grained matching performance. It has a better robustness in handling highly similar hard negatives, which provides a new way to address cross-modal hard sample discrimination.