Skip to content
Conference

Dual Associations Semantic Enhancement for Image-Text Matching

2026 · Poster Volume 0008 The 2026 Twenty-Second International Conference on Intelligent Computing July 23-26, 2026 Toronto, Canada · 0 citations

Abstract

Image-text matching faces the significant challenge in effectively mitigating the large visual-semantic discrepancy among different modalities. Existing studies mainly address this by projecting multi-modal data into a common subspace for measuring their semantic similarities. However, most of them often overlook the inherent asymmetry for directional associations, i.e., the differences between im-age-to-text and text-to-image affinities, which often results in inaccurate retrieval results. To tackle the challenge, we propose a Dual Associations Semantic En-hancement (DASE) model to capture bidirectional image-text semantic associa-tions. Specifically, we first build a two-layer GCN fusion network to construct and mine the semantic associations for each modality. And then, due to the inher-ent asymmetry of directional associations, a Dual Associations Alignment Mod-ule (DAAM) is designed to capture the dual associations between visual and tex-tual modalities, enabling comprehensive cross-modal fine-grained interaction. Fi-nally, global alignment is incorporated with the local alignment to achieve full semantic matching across heterogeneous modalities in a unified embedding space. Experimental results on two publicly available datasets demonstrate that the pro-posed DASE model achieves significant performance improvements in image-text matching tasks compared to baseline methods, validating its effectiveness and superiority.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.