Jul 2026· IEEE Transactions on Image Processing· Vol 35, pp. 7521-7534· 0 citations· 51 references
MedicineComputer Science
Abstract
Unsupervised person re-identification (ReID) aims to learn identity-discriminative representations without manual annotations, which is challenging due to noisy pseudo labels, background clutter, and large appearance variations. Recent studies have shown that exploiting fine-grained local cues is crucial for improving robustness in unsupervised ReID. In this context, random masking has emerged as a simple and annotation-free way to encourage the model to focus on informative regions. However, existing masking-based unsupervised ReID methods still suffer from two limitations: (1) Underused masked views: masked views are treated as degraded auxiliaries rather than exploited as fine-grained supervisory signals; (2) Weak cross-view alignment: feature alignment is restricted to mini-batch pairs, lacking explicit global alignment between masked and unmasked views across clusters. To address these issues, we propose the Mask-guided Asymmetric Contrastive and Semantic Alignment (ACSA) framework. Specifically, we introduce an Asymmetric Contrastive Learning (ACL) module with a dual-memory mechanism to separately encode masked and unmasked features, allowing masked views to serve as informative and discriminative supervision. In parallel, a Semantic Alignment Learning (SAL) module conducts multi-granularity distribution alignment by aligning both cluster-level prototypes and randomly sampled instance-level features, thereby preserving semantic consistency and intra-cluster diversity. Furthermore, to provide more reliable semantic anchors for SAL under noisy pseudo labels, we introduce a Progressive Refinement Module (PRM), which refines prototypes and features via exponential moving averaging for more stable semantic alignment. Extensive experiments validate the superiority of our method, even outperforming certain supervised counterparts. Code is available at https://github.com/Trangle12/ACSA
Experimental results on multiple benchmark datasets demonstrate that the UGCL framework exhibits favorable robustness and competitive performance under both noisy and noise-free correspondence conditions.
Wen Ting, Shidu Dong, Yuzhi Zhang et al.· Machine Vision and Applicati...· 0 citations
Unsupervised person re-identification (USL-ReID) typically relies on clustering to generate pseudo-labels, but significant cross-view appearance variations often cause images of the same identity to be split into different clusters. Training on such noisy pseudo-labels severely degrades the learned representations. Therefore, learning robust view-invariant features is paramount. Data augmentation provides a direct way to enhance invariance, yet its trade-offs in USL-ReID remain under-explored: weak augmentations usually preserve identity semantics but lack diversity, whereas strong augmentations provide richer appearance diversity at the cost of partially corrupting identity-consistent semantic cues. To address this challenge, we propose Invariant Representation learning with Progressive Prototype Refinement (IRPP), a unified framework that learns invariant and discriminative features from noisy pseudo-labels. IRPP consists of three synergistic components. First, an Augmented Dual-Contrastive Learning (ADCL) module performs dataset-level prototype-guided invariant learning by contrasting weakly and strongly augmented views against cluster-derived prototypes. Second, an Alignment and Uniformity Learning (AUL) module regularizes the mini-batch-level weak–strong feature geometry, leading to more stable feature distributions under data augmentation. Third, a Progressive Prototype Refinement (PPR) mechanism progressively optimizes cluster centroids into cleaner prototypes, thereby mitigating the influence of noisy pseudo-labels and further strengthening invariant representation learning. This closed-loop design enables prototype-guided contrastive learning, weak–strong regularization, and prototype refinement to mutually reinforce each other. Extensive experiments on standard USL-ReID benchmarks demonstrate that IRPP achieves state-of-the-art performance with a simple and efficient training pipeline. Code is available at https://github.com/Trangle12/IRPP
Xuan Tan, Qixian Zhang, Ding Qi et al.· IEEE Transactions on Image P...· 0 citations
Structural-Semantic Reciprocal Learning (SSRL), a framework that transforms open-loop association into a self-correcting closed-loop system, achieves robust cross-modal representation through the reciprocal interaction between structural and semantic learning.
Moyao Tian, Shijia Liu, Yan Yang et al.· arXiv.org· 0 citations
This work proposes CLIP-SGI, a semantic-guided and instance-consistent framework for generalizable person ReID that combines semantic guidance, domain-aware representation learning, and instance consistency to improve robustness under domain shifts.
Dai-Xin Liu, Yu Yang, Linlin Tang et al.· IEEE Transactions on Image P...· 0 citations
Text-Based Person Retrieval (TBPR) aims to locate a person in an image database based on a natural language description. While effective in theory, TBPR faces substantial challenges in real-world scenarios due to noisy correspondences—misaligned or weakly related image-text pairs—that significantly degrade retrieval performance. Existing methods often overemphasize hard negative mining, which inadvertently magnifies the impact of such noise. To address this issue, we propose Dynamic Uncertainty with Noisy Correspondences (DUNC), a novel framework that incorporates two key components: (1) Cross-modal Evidential Learning (CEL), which models bidirectional alignment uncertainty using a Dirichlet distribution to capture the confidence in image-text similarity, and (2) Dynamic Robust Loss (DRL), which adaptively selects and aggregates hard negative samples to reduce the influence of noisy instances and improve model robustness. Unlike conventional global-alignment approaches, DUNC exploits fine-grained local correspondences to enhance semantic alignment between modalities. By integrating uncertainty-aware modeling and adaptive contrastive supervision, our method is capable of effectively disentangling noisy from reliable training pairs. Extensive experiments conducted on three benchmark datasets—CUHK-PEDES, ICFG-PEDES, and RSTPReid—demonstrate that DUNC consistently achieves state-of-the-art performance and exhibits strong robustness across a wide range of noise conditions. Code is publicly available at https://github.com/ASL-forever/DUNC.
Zequn Xie, Chuxin Wang, Sihang Cai et al.· ACM Transactions on Informat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.