A cross-modal fusion detection framework for robust object perception under challenging illumination
Abstract
RGB-infrared (RGB-IR) fusion improves object detection by enabling robust object localization under challenging illumination and environmental conditions. However, the additional IR modality often increases computational cost, limiting its deployment in real-time measurement and perception systems. This work revisits cross-modal fusion strategies and shows that strong performance can be achieved by progressively accumulating efficient cross-modal interactions across multiple feature scales, without duplicating auxiliary modality feature-extraction branches. To exploit this observation, a cross-modal fusion (CMFusion) detection framework is proposed, which focuses on efficient multi-level RGB-IR feature interaction within YOLO-based detectors. This design reduces redundant feature extraction within individual modalities while enhancing multi-level multimodal interaction, enabling effective cross-modal complementarity with lower computational overhead. Extensive experiments on three public RGB-IR datasets across two detection tasks (OBB and HBB) demonstrate that CMFusion achieves a 1.7%–6.1% mAP improvement over baseline unimodal detectors while maintaining comparable low computational complexity, and consistently surpasses state-of-the-art methods.