Skip to content

Mitigating Noisy Correspondence in Video-Text Retrieval via Noise-mined Adaptive Self-Labeling

Jul 2026 · ACM Transactions on Multimedia Computing, Communications, and Applications (TOMCCAP) · 0 citations · 58 references

Abstract

In video-text retrieval, addressing the noisy correspondence problem is crucial for achieving accurate retrieval performance. Recent methods address this challenge by either suppressing the impact of noisy data or predictions distorted by noise. However, existing approaches often overlook two critical aspects: distinguishing hard noise from semantically ambiguous cases, and preserving latent associations within unmatched negative pairs. To address this limitation, we propose Noise-mined Adaptive Self-Labeling (NASL), which effectively manages noisy data during training. NASL consists of two loss functions: 1) Noise-mined Matching Loss (NML), which identifies and penalizes noisy data based on a two-stage suppression strategy, and 2) Adaptive Self-labeling Loss (ASL), which employs optimal transport to recover latent associations among false negatives in noisy conditions and provides soft supervision to prevent excessive penalization of semantically plausible pairs. Extensive experiments demonstrate that NASL improves the separation between clean and noisy data while effectively mitigating noise, leading to significant performance improvements on the MSR-VTT, DiDeMo, MSVD, LSMDC, and ActivityNet datasets.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.