Multi Scale Feature Alignment Network for Cross-Modal Visible-Infrared Person Re-Identification
To address the performance degradation of person re-identification (ReID) under complex lighting and day-night conditions, this study proposes a novel dual-path convolution based multi-scale feature alignment (DCMFA) network. The network mainly focuses on addressing the challenges of modality discrepancy and feature alignment between visible and infrared images. First, to cope with feature deformation in persons caused by external factors in both modalities, we design a dual-path convolution (DPCon) layer. Secondly, to further alleviate feature loss caused by scale variation, we construct a multi-scale feature aggregation (MSFA) module by stacking DPCon layers of different depths and incorporating an attention mechanism to effectively aggregate key multi-scale information and suppress redundancy. Finally, to enable more effective alignment of multi-scale features across the two modalities, we propose a feature mapping and alignment operations (FMAO) module. Experimental results on three publicly available cross-modal visible-infrared person re-identification (VI-ReID) datasets demonstrate that our DCMFA network significantly outperforms existing mainstream methods in terms of recognition accuracy. Specifically, on the SYSU-MM01 dataset, our method achieves a Rank-1 accuracy of 83.74% in the single-shot setting and 89.47% in the multi-shot setting.