A small and dense target detection framework based on feature perception and adaptive fusion
Abstract
Small and dense object detection remains challenging in complex visual scenes. Repeated downsampling weakens discriminative features of tiny objects, while dense object distributions cause severe feature overlap and semantic ambiguity. To address these challenges, this paper proposes Enhanced Feature-Aware YOLO (EFA-YOLO), a lightweight detection framework based on YOLO11n. The framework establishes a collaborative optimization architecture in which feature extraction, cross-scale feature fusion, and small-object localization are jointly optimized through three tightly coupled components. The Multi-Scale Frequency-Aware Attention module uses Haar wavelet as a basic frequency decomposition tool and integrates multi-scale feature extraction, multi-directional high-frequency enhancement, frequency-aware spatial attention, and semantic-guided feature calibration to preserve discriminative information for small and dense objects. The Adaptive Weighted Concatenation Fusion module introduces learnable weighted concatenation, instead of direct weighted feature aggregation, to enhance bidirectional cross-scale feature interaction while preserving fine-grained feature information. In addition, a lightweight high-resolution detection head combines adaptive cross-layer feature fusion with redundant branch pruning to improve localization accuracy for tiny objects with limited computational overhead. Experiments on the VisDrone and COCO datasets show that, compared with YOLO11n, EFA-YOLO improves mAP50 by 3.7% and 2.8%, respectively. Ablation studies further verify the effectiveness of each component and their collaborative contributions to overall detection performance. The proposed framework achieves a favorable balance between detection accuracy and computational efficiency, demonstrating real-time performance on the evaluated GPU platform and indicating potential for further edge-oriented optimization.