ISGM: An Illumination-Aware Semantic-Guided Mamba Network for RGB–Infrared Vehicle Detection in UAV Imagery
Abstract
RGB–infrared (RGB-IR) vehicle detection in uncrewed aerial vehicle (UAV) imagery is essential for applications, such as traffic monitoring and object tracking. However, existing methods often suffer from heterogeneous feature responses across modalities, degraded RGB feature representations under adverse illumination, and insufficient capture of fine-grained structural and edge details, which collectively impede accurate cross-modal modeling and localization. To alleviate these issues, we propose an Illumination-aware Semantic-Guided Mamba (ISGM) network. First, we design a Feature Representation Refinement module that stabilizes RGB and IR features through scale-specific channel remapping and normalized nonlinear refinement, yielding more reliable representations for subsequent cross-modal interaction. Furthermore, to enhance robustness to illumination variations and better preserve detailed structural cues and boundary information, we develop a cross-modal feature interaction mechanism comprising the Illumination-Aware Fusion Modulation (IAFM) module and the Detail-enhanced Semantic-Guided Mamba (DSGM) module. Specifically, the IAFM module estimates illumination-aware modality reliability weight maps, thereby improving robustness under challenging illumination conditions. These weight maps guide the DSGM module to integrate high-level semantic information and low-level detail cues into multiscale guidance features. These features are subsequently used to modulate the scanning parameters for adaptive RGB-IR feature interaction. This design improves the modeling of target-region features while minimizing interference from complex backgrounds. Extensive experiments on the DroneVehicle and VEDAI datasets demonstrate that ISGM outperforms state-of-the-art methods in detection performance while achieving a favorable accuracy-efficiency balance among comparable methods.