Frequency-Guided Cross-Scale Refinement Network for UAV Detection
Abstract
In recent years, the use of UAVs has become increasingly widespread, and the public safety risks posed by unauthorized UAV flights have become increasingly prominent, creating an urgent need for effective detection and identification of UAV targets. However, such targets are small in size, have low contrast, and exhibit an extremely low signal-to-noise ratio; conventional detection methods generally suffer from insufficient feature discrimination, missed detections, and false alarms in complex backgrounds. To address these challenges, this paper proposes a Frequency-Guided Cross-scale Refinement Network (FGCR-Net). Based on an encoder-decoder architecture, this network achieves end-to-end collaborative optimization through cross-layer feature fusion, side-channel prediction refinement, and frequency-domain background suppression. First, a multi-path selective cross-layer fusion module (SCFM) is designed. This module employs coordinated modeling via both channel and spatial paths, supplemented by adaptive weighting with learnable coefficients, to perform differentiated selective fusion of the encoder’s fine-grained features and the decoder’s semantic features, thereby bridging the semantic gap at jump connections; Second, we designed a Cross-Scale Adaptive Fusion Enhancement Attention Module (CAFEM), which cascades multi-receptive-field hollow convolutions, strip pooling, and a bidirectional semantic guidance mechanism to perform cross-scale refinement on the side outputs of each decoder layer, thereby alleviating the issues of blurred boundaries and false alarms caused by inconsistent quality of multi-scale prediction maps and insufficient cross-layer consistency; finally, we design a Frequency-Guided Semantic Enhancement Module (FGSEM), which uses the Fast Fourier Transform (FFT) to decouple encoder features into the frequency domain. By leveraging low-frequency energy to predict the background confidence map and applying spatially selective suppression to high-frequency components, this module distinguishes, from a frequency-domain perspective, the high-frequency responses of complex backgrounds and targets that are highly similar in the spatial domain. Experiments on MSDS-UAV, a self-built multi-scenario UAV dataset for small targets, demonstrate that our method consistently outperforms existing state-of-the-art methods across multiple performance metrics, with Pixel Accuracy, Mean Intersection over Union, and Probability of Detection reaching 92.76%, 70.91%, and 92.69%, respectively; Compared to the baseline model, these three metrics improved by 1.90, 3.20, and 3.76 percentage points, respectively, fully validating the effectiveness and superiority of the proposed method.