Sep 2026· International Journal of Machine Learning and Cybernetics· Vol 17· 0 citations· 58 references
TL;DR
A novel Robust Modality Unified Learning framework (RMUL), consisting of a Robust Cross-modality Label Transfer (RCLT) method and Modality Unified Learning (MUL) module, which unifies cross-modality labels only for modality-shared instances while retaining the original labels for modality-specific ones.
Unsupervised visible-infrared person re-identification (USL-VI-ReID) learns person representations that can be compared across modalities without identity annotations. In the unpaired setting, however, identity correspondences between modalities are often incomplete, leaving many identities without an observed counterp...
Duan-Ning Chen, Ke He, Bin Yang et al.· 0 citations
A Progressively Biased Split Vision Transformer (PBSVT) is proposed, which combines a split ViT backbone with progressive bias training to gradually reduce RGB-dominant bias while preserving modality-shared structure and demonstrates the effectiveness of progressive modality transition for robust VI-ReID representation...
Mengru Jiao, Xin-Yue Xu, Jun-Feng Zhang· International journal of pat...· 0 citations
It is argued that VI-ReID should be treated as an early cross-modal correspondence learning problem rather than only a late embedding alignment problem, and CMIA-Net is proposed, a framework that establishes bidirectional visible-infrared interaction at shallow backbone stages and introduces Spectral-Invariant Augmenta...
Dao-Li Zhang, Qi-Cheng Liu· Engineering Research Express· 0 citations
Unsupervised person re-identification (USL-ReID) typically relies on clustering to generate pseudo-labels, but significant cross-view appearance variations often cause images of the same identity to be split into different clusters. Training on such noisy pseudo-labels severely degrades the learned representations. The...
Xuan Tan, Qi-Xian Zhang, Ding Qi et al.· IEEE Transactions on Image P...· 0 citations
A unified empirical evaluation on the Market-1501, Occluded-DukeMTMC, and MSMT17 datasets is presented, revealing that latent feature-space refinement and semantic cross-modal alignment offer superior stability, scalability, and robustness compared to explicit pixel-level generation.