Structural-Semantic Reciprocal Learning (SSRL), a framework that transforms open-loop association into a self-correcting closed-loop system, achieves robust cross-modal representation through the reciprocal interaction between structural and semantic learning.
Abstract
Unsupervised visible-infrared person re-identification (USVI-ReID) is challenging due to the large modality gap and the lack of cross-modal identity annotations. Progressive association paradigms have been proposed to gradually bridge the gap, but they suffer from two critical bottlenecks: reliance on ambiguous global representations and unchecked propagation of pseudo-label noise in an open-loop manner. To address these issues, we propose Structural-Semantic Reciprocal Learning (SSRL), a framework that transforms open-loop association into a self-correcting closed-loop system. Structurally, we introduce Fine-grained Structural Decoupling (FSD) to extract discriminative body-part primitives as reliable spatial anchors, complementing ambiguous holistic silhouettes with spatially consistent structural details. Semantically, we design a Closed-loop Semantic Calibration (CSC) mechanism that reconstructs shared semantic prototypes at each epoch and feeds them back into the training loop, effectively filtering pseudo-label noise before the next clustering cycle. Through the reciprocal interaction between structural and semantic learning, SSRL achieves robust cross-modal representation. Extensive experiments demonstrate the competitive performance of SSRL against state-of-the-art USVI-ReID methods on both SYSU-MM01 and RegDB, notably surpassing several supervised counterparts on RegDB.
Unsupervised person re-identification (USL-ReID) typically relies on clustering to generate pseudo-labels, but significant cross-view appearance variations often cause images of the same identity to be split into different clusters. Training on such noisy pseudo-labels severely degrades the learned representations. Therefore, learning robust view-invariant features is paramount. Data augmentation provides a direct way to enhance invariance, yet its trade-offs in USL-ReID remain under-explored: weak augmentations usually preserve identity semantics but lack diversity, whereas strong augmentations provide richer appearance diversity at the cost of partially corrupting identity-consistent semantic cues. To address this challenge, we propose Invariant Representation learning with Progressive Prototype Refinement (IRPP), a unified framework that learns invariant and discriminative features from noisy pseudo-labels. IRPP consists of three synergistic components. First, an Augmented Dual-Contrastive Learning (ADCL) module performs dataset-level prototype-guided invariant learning by contrasting weakly and strongly augmented views against cluster-derived prototypes. Second, an Alignment and Uniformity Learning (AUL) module regularizes the mini-batch-level weak–strong feature geometry, leading to more stable feature distributions under data augmentation. Third, a Progressive Prototype Refinement (PPR) mechanism progressively optimizes cluster centroids into cleaner prototypes, thereby mitigating the influence of noisy pseudo-labels and further strengthening invariant representation learning. This closed-loop design enables prototype-guided contrastive learning, weak–strong regularization, and prototype refinement to mutually reinforce each other. Extensive experiments on standard USL-ReID benchmarks demonstrate that IRPP achieves state-of-the-art performance with a simple and efficient training pipeline. Code is available at https://github.com/Trangle12/IRPP
Xuan Tan, Qixian Zhang, Ding Qi et al.· IEEE Transactions on Image P...· 0 citations
A disentanglement method based on Mutual Information Minimization is introduced to minimize statistical dependence between modality-shared and modality-specific features from a probability distribution perspective and a Local Discriminative Attention Module is designed to adaptively focus on highly informative body parts such as head-shoulder ratio and torso patterns.
Xiaokai Liu, Fangqing Zhou, Qian Song et al.· Advances in Engineering Tech...· 0 citations
Visible-infrared person re-identification (VI-ReID) is an important technique for around-the-clock person matching, and its primary challenge arises from substantial cross-modal discrepancies. To address this challenge, we propose a Decoupled Information-Guided Cross-Modal Alignment (DIGCA) framework organized into three stages: multi-scale feature modeling, partial functional decoupling, and guided cross-modal alignment. First, a Hierarchical Context Extractor (HCE) aggregates multi-scale contextual information through dilated convolutions and residual connections to enrich the initial identity representation. Second, a Multi-Branch Unified Encoder (MBUE) organizes complementary feature streams through parallel global-relation and local-spatial modeling. The resulting representations exhibit different information tendencies and promote partial functional decoupling between identity-related and modality-related information. Finally, the Decoupled Information-Guided Cross-Modal Alignment (DIG-CMA) module refines the two streams with channel and spatial attention and integrates them through cross-attention. Under the joint optimization objective, the resulting representations support cross-modal alignment. Experiments on SYSU-MM01, RegDB, and LLCM show that DIGCA achieves competitive overall performance compared with recent VI-ReID methods, providing empirical support for the effectiveness of the proposed decoupling-guided alignment strategy in cross-modal identity matching.
Unsupervised Visible-Infrared Person Re-identification (US-VI-ReID) aims to achieve cross-modal identity matching without manual annotations. However, existing methods overlook the unreliable cross-modal cluster associations caused by hard matching and pseudo-label noise generated by cross-modal joint clustering. To address these issues, this paper proposes a novel Distribution-Aware Soft Alignment with Boundary Correction (DASBC) framework which consists of two collaborative modules, the Distribution-Aware Soft Alignment (DASA) module and the Boundary Pseudo-Label Correction (BPLC) module. The DASA module constructs cross-modal inter-cluster associations via a probabilistic soft alignment mechanism, quantifying the association strength between cluster pairs to flexibly and robustly capture both strong associations and weak ambiguous relationships. The BPLC module leverages the Silhouette Score to locate cluster boundary samples, verifies and corrects label rationality to eliminate noise, and provides reliable supervision signals. In addition, this paper designs a multi-level contrastive learning loss that integrates unimodal contrast, cross-modal alignment, and unified contrast supervision to learn modality-invariant feature representations. Extensive experiments on the SYSU-MM01 and RegDB datasets demonstrate the superiority of the proposed DASBC framework.
Hongyang Fu, Jin Wang, Xiao-Shuai Niu et al.· Neural Networks· 0 citations
Unsupervised vehicle re-identification (Re-ID) is pivotal for scalable intelligent transportation systems but faces significant challenges from severe noise accumulation. Traditional clustering-based methods often suffer from error propagation during online training, as purely visual features are highly susceptible to intra-class viewpoint variations and inter-class similarities. To mitigate this fundamental limitation, we propose a novel Semantically-Anchored Dynamics (SAD-ReID) framework that exploits the viewpoint-invariant stability of text-induced semantic knowledge derived from Vision-Language Models. Specifically, we first introduce a Semantically-Guided Initialization strategy that fuses visual similarities with detailed textual descriptions generated automatically by Qwen-VL. This rectifies initial visual clusters to establish robust text-guided dual-prototype (visual and semantic) anchors. During the online learning phase, we propose a Reliability-Aware Dynamic Update (RADU) mechanism. By calculating a Cross-Modality Allegiance (CMA) score that measures the topological agreement between visual and textual spaces, RADU dynamically adjusts the memory update momentum. This efficiently accelerates learning from reliable samples while filtering out noisy pseudo-labels to prevent memory corruption. Furthermore, an Adaptive Granularity Attention Fusion (AGAF) module is designed to capture both global semantic attributes and fine-grained local discriminative details. Extensive experiments on the VeRi-776 and VehicleID benchmarks demonstrate the significant superiority of our approach over existing state-of-the-art methods, achieving an impressive Rank-1 accuracy of 90.0% and 43.2% mAP on VeRi-776. By enforcing text-visual semantic consistency throughout the evolutionary training process, SAD-ReID successfully prevents model drift and learns highly robust representations without any manual annotation.
Yun Jiang, Kunyi Zhu, Tao Sun· Journal of King Saud Univers...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.