Small-Target Detection via Fusion of Visible and Infrared Image Features
Abstract
Visible–infrared small-target detection is challenged by weak single-modality representation, modality discrepancy, and the quadratic cost of dense cross-modal attention. We propose TFFB, a feature-level fusion detector that combines spatial feature compression (SFC), cross-attention modality enhancement (CME), and iterative cross-modal enhancement (ICME) to balance information exchange and computational efficiency. To further improve localization, we introduce Focaler-SIoU for small-box regression. On Anti-UAV300, TFFB improves the middle-fusion baseline from 76.2%/43.7% to 81.5%/48.6% in mAP@0.5/mAP@0.5:0.95, and TFFB with Focaler-SIoU reaches 83.6% and 50.2%, respectively. The results indicate that compact cross-modal interaction can strengthen visible–infrared UAV detection while keeping computational costs moderate.