A unified empirical evaluation on the Market-1501, Occluded-DukeMTMC, and MSMT17 datasets is presented, revealing that latent feature-space refinement and semantic cross-modal alignment offer superior stability, scalability, and robustness compared to explicit pixel-level generation.
A novel Robust Modality Unified Learning framework (RMUL), consisting of a Robust Cross-modality Label Transfer (RCLT) method and Modality Unified Learning (MUL) module, which unifies cross-modality labels only for modality-shared instances while retaining the original labels for modality-specific ones.
Zhiyong Li, Wei Jiang, Haojie Liu et al.· International Journal of Mac...· 0 citations
Experiments show that DIGCA achieves competitive overall performance compared with recent VI-ReID methods, providing empirical support for the effectiveness of the proposed decoupling-guided alignment strategy in cross-modal identity matching.
A novel end-to-end Diffusion-Driven Dual-stream Framework (D 3 F), which seamlessly integrates generative structural priors from Diffusion Transformers (DiT) into vision-language ReID, achieving state-of-the-art (SOTA) performance on both occluded and holistic ReID benchmark datasets.
It is argued that VI-ReID should be treated as an early cross-modal correspondence learning problem rather than only a late embedding alignment problem, and CMIA-Net is proposed, a framework that establishes bidirectional visible-infrared interaction at shallow backbone stages and introduces Spectral-Invariant Augmenta...
Dao-Li Zhang, Qi-Cheng Liu· Engineering Research Express· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.