ZOS-Net: A Lightweight RGB-T Object Detection Network with Cross-Modal Relation Enhancement and Selective Target Awareness
Abstract
In tasks such as autonomous driving, low-altitude remote sensing, intelligent surveillance, and low-altitude security, RGB-T object detection must maintain stable performance under complex conditions, including illumination variations, thermal-source interference, background clutter, and dense distributions of small objects. Existing multimodal detection methods typically rely on complex attention structures or heavyweight fusion modules, making it difficult to balance detection accuracy, model lightweightness, and deployment efficiency; moreover, modality noise and redundant background information are prone to joint propagation during shallow fusion, weakening the responses of small and weak objects. To address these issues, this paper proposes ZOS-Net, a lightweight RGB-T object detection network that integrates cross-modal relation enhancement and selective target awareness. Specifically, ZOS-Net introduces a Cross-Modal Fine-Grained Gated Fusion module, termed ZCGF, at the shallow P3 stage to enhance reliable complementary information from visible textures and infrared thermal responses; an Object-Aware Relation Enhancement module, termed OAGR, is introduced at the semantic bridging stage from P4 to P3 to generate object-relation priors; and a Selective Relation-Guided Target Perception Enhancement Module, termed SRTPEM, is further employed to perform foreground enhancement and weak background calibration on the final P3 features. Experimental results show that ZOS-Net achieves a favorable balance between detection performance and complexity on the M3FD, FLIR-aligned, and VEDAI datasets. On M3FD, ZOS-Net achieves 88.3% mAP50 and 59.5% mAP50:95 with only 4.74 M parameters and 7.9 G FLOPs; on FLIR-aligned, it achieves 83.5% mAP50 and 46.8% mAP50:95; and on VEDAI, it achieves 74.5% mAP50. These results indicate that the proposed method improves the detection performance of small and weak objects in complex backgrounds and provides a lightweight solution for RGB-T object detection on resource-constrained platforms.