Aug 2026· Applied Sciences· 0 citations· 31 references
TL;DR
HD-YOLO improves small-object detection with a compact parameter footprint, while direct hardware benchmarks remain necessary to establish deployment efficiency.
Abstract
Unmanned aerial vehicle (UAV) imagery supports intelligent surveillance, environmental monitoring, traffic management, and infrastructure inspection. Yet aerial detection is difficult when objects are small, crowded, and observed at markedly different scales. Background clutter and illumination changes further weaken target cues and impair localization. We therefore propose HD-YOLO, a lightweight multiscale detector for small objects in UAV imagery. Its Multi-Dilation Shared Convolution Kernel (DSCK) extracts local texture and contextual information with shared dilated kernels. The Hybrid Dilated Bidirectional Feature Pyramid Network (HDFPN) reconstructs global and local cues before bidirectional aggregation, enabling high-resolution evidence to reach the prediction layers. The Efficient and Slim Head (ES-Head) combines shared operations with differential convolution to reduce cost and strengthen boundary-sensitive features. A joint ShapeIoU and Normalized Wasserstein Distance loss improves regression for small, irregular objects. Together, these components reduce missed detections in dense, cluttered scenes without relying on large model capacity. On VisDrone2019, HD-YOLO improves precision, recall, mAP50, and mAP50:95 over YOLOv8n by 6.9%, 7.2%, 8.2%, and 5.2%, respectively, while reducing parameters from 3.0 M to 0.9 M. Evaluations on TinyPerson and HIT-UAV also support its utility for tiny pedestrians and infrared aerial targets. HD-YOLO therefore improves small-object detection with a compact parameter footprint, while direct hardware benchmarks remain necessary to establish deployment efficiency.
A Lightweight Feature-Fusion and Small-Target Enhancement Network (LFE-YOLO), a lightweight detector that coordinates partial-channel feature extraction, efficient cross-scale fusion, high-resolution prediction, background-interference suppression, and stable tiny-box regression within a unified architecture is proposed.
Mingxi Chen, Cheng Guo, Shao-Jie Ma et al.· Drones· 0 citations
Unmanned aerial vehicle (UAV) imagery is widely used in urban monitoring, public security, and disaster assessment. However, object detection in UAV scenes faces multiple challenges, including a high proportion of small objects, severe occlusion in crowded areas, complex background textures, and image degradations such as haze, which often cause generic detectors to suffer from missed detections, false alarms, and unstable localization. To address these issues, we propose a lightweight multi-scale enhanced detection framework tailored for complex UAV scenarios. Built upon a MobileNet backbone, the proposed framework introduces a multi-scale enhancement module that constructs a feature pyramid and incorporates a scale-adaptive fusion mechanism to dynamically reweight the contributions of features from different scales. In addition, a fine-grained detail enhancement branch is deployed at high-resolution levels to strengthen edge and texture cues, while a context compensation module is designed to alleviate local uncertainty under dense occlusion and low-contrast conditions, thereby improving small-object separability and localization stability. Experimental results demonstrate that our method achieves strong performance on both VisDrone-DET and HazyDet, reaching mAP@0.5 = 0.312 and mAP@0.5:0.9 = 0.167 on VisDrone-DET, and mAP@0.5 = 0.719 and mAP@0.5:0.9 = 0.483 on HazyDet. The proposed method also shows favorable efficiency on an RTX 4060 Ti desktop GPU, indicating its real-time inference potential under the reported desktop hardware setting.
Xuehua Tao, Ji-Wei Sun· Engineering Research Express· 0 citations
A context-gated dynamic perception framework that treats small-object feature degradation as a coupled problem of representation, fusion, and prediction and indicates a practical accuracy-efficiency trade-off for dense aerial small-object perception.
Guangjun Gao, Ruibing Xie· Pattern Analysis and Applica...· 0 citations
Vehicle detection from heterogeneous traffic imagery is essential for large-scale traffic monitoring. However, satellite remote sensing, uncrewed aerial vehicle (UAV), and closed-circuit television (CCTV) images differ substantially in spatial resolution, viewing geometry, illumination conditions, and vehicle scale, making unified cross-source detection challenging, especially for small targets and unseen source-domain distributions. To address these issues, this study proposes a lightweight source-conditioned multiscale detection framework for vehicle detection from independent satellite, UAV, and CCTV image domains. The framework is centered on a source-aware multiscale awareness fusion module (SA-MS-AFM), which performs source-expert reweighting from a single input image and supports masked-expert training for incomplete expert availability. SwiftFormer with efficient additive attention (EAA) is adopted for efficient local–global representation, while Shiftwise convolution enhances local structural cues with low overhead. Content-aware reassembly of features (CARAFEs) is introduced in the neck to reduce spatial-detail loss during top-down feature reconstruction, and a four-scale prediction head is designed to accommodate large vehicle-scale variations. Experiments on a pooled cross-source dataset of 6805 images show that the proposed model achieves an $F_{1}$ -score of 0.885, mAP@0.5 of 0.886, and mAP@0.5:0.95 of 0.513 while maintaining 41.53 frames per second (FPS) on the tested RTX 4090 GPU. Ablation, expert-masking, three-scale/four-scale tradeoff, and zero-shot external generalization tests further demonstrate the effectiveness and transferability of the proposed framework under heterogeneous traffic imaging conditions.
Zhen Liu, Mei-Po Kwan, Weiwei Jiang et al.· IEEE Transactions on Geoscie...· 2 citations
Unmanned Aerial Vehicle (UAV) aerial photography is extensively utilized in security, traffic monitoring, and disaster rescue. However, UAV-captured images present significant challenges, including small target scales, dense distribution, and complex backgrounds. While conventional object detection algorithms like the YOLO series have made progress, they often struggle to balance accuracy and real-time performance in these resource-constrained environments. To address these issues, we propose SSM-YOLO11s, a lightweight model optimized for small object detection in aerial imagery. Our approach first introduces the Sitou module, which employs a deep-channel compression and shallow feature retention strategy with a secondary fusion branch to reduce parameters by 50% while enhancing fine-grained feature utilization. Furthermore, the lightweight SNGSConvE module is designed by integrating SNI, GSConvE, and CSPOmniKernel to mitigate feature misalignment and strengthen capture capabilities. Finally, a Multi-Scale Edge Enhancement (MSEE) module is constructed to fuse edge details across multiple scales, improving target discriminability. Experimental results on the VisDrone2019 dataset demonstrate that SSM-YOLO11s achieves a superior balance between precision and efficiency compared to state-of-the-art models.
Junfu Chen, Xi Zhao· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.