Skip to content

Author

Zedong Huang

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

A frequency-saliency guided multi-scale feature fusion network for robust UAV object detection

Unmanned aerial vehicle (UAV)-based object detection plays an important role in intelligent surveillance and traffic monitoring. However, small object scales, complex backgrounds, dense distributions, and occlusions make it difficult to balance detection accuracy and real-time performance. To address these challenges, this paper proposes a frequency-saliency guided multi-scale feature fusion network based on YOLO11 for UAV object detection. Specifically, a Global Frequency Feature Extraction module is designed by integrating saliency enhancement and Dual-Tree Complex Wavelet Transform to separate high-frequency target details from low-frequency background information. In addition, a lightweight multi-scale feature fusion module with soft gating and spatial attention is introduced to enhance feature representation across scales, and a Scale-Adaptive Dropout Feature Pyramid Network is employed to enhance robust multi-scale feature learning. Experimental results on the HIT-UAV dataset show that FSMF-YOLO achieves 91.6% precision, 83.5% recall, and 89.3% mAP@50 at 101 FPS, improving mAP@50 by 2.6% over YOLO11n while maintaining real-time inference. These results suggest that frequency-saliency guidance improves small-object representation in complex UAV scenes and provides a favorable accuracy-efficiency trade-off compared with lightweight YOLO-based detectors.

Linji Cheng, Fan Guo, Zedong Huang · 0 citations
Open access Jul 2026

Object detection algorithm based on infrared-visible dual-modality feature fusion

CFM-YOLO, an infrared–visible dual-modality detection algorithm based on YOLOv11, is proposed to improve UAV object detection under adverse illumination and complex aerial backgrounds. The network is redesigned from three aspects: cross-modal feature extraction, lightweight feature fusion, and small-object-oriented detection. First, a Cross-Modality Fusion Mamba (CFM) module is introduced to promote channel-level interaction between visible and infrared features and to model long-range spatial dependencies with selective state-space modeling. Second, a lightweight feature fusion network is used to improve multi-scale information transmission while limiting redundant computation. Third, a P2 detection layer, Ghost convolution, and Focal-WIoU loss are incorporated to enhance small-object localization and alleviate the effect of imbalanced bounding-box samples. Quantitative experiments on the DroneVehicle dataset show that CFM-YOLO achieves 81.6% mAP@0.5 and 60.3% mAP@0.5:0.95, improving over the YOLOv11n-dual/base baseline by 8.4 and 5.5 percentage points, respectively. Qualitative results on the LLVIP dataset further indicate that the proposed method can reduce several missed detections in low-light pedestrian scenes. These results suggest that CFM-YOLO provides a competitive trade-off between detection accuracy and computational cost for UAV-based infrared–visible object detection.

Zedong Huang, Kang-Kang Du, Xiao-Huang Hu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.