Tiny object detection in remote sensing typically faces the challenges of being submerged in backgrounds, limited feature representation, and high sensitivity to prediction errors due to the small size and diverse shapes. To address these challenges, a geometric guided noise reduction super-resolution (SR) network is proposed. First, a dual-branch denoising and SR feature pyramid network is proposed, which integrates an adaptive dynamic noise reduction module and an inference decoupled auxiliary SR branch, while a progressive loss-annealing strategy is further introduced to reduce reliance on the SR branch during inference, meeting the requirements of lightweight and high-performance remote sensing tiny object detection. Second, a geometric characteristic regression metric is proposed, which comprehensively considers the location accuracy and the shape similarity between the prediction and ground truth boxes, thereby improving bounding-box quality and detection precision. Extensive experiments have been conducted on the remote sensing tiny object datasets AI-TOD v1, AI-TOD v2, USOD, and VisDrone. Specifically, it reaches an AP of 31.6 on AI-TOD v1, 30.5 on AI-TOD v2, 37.4 on USOD, and 30.5 on the VisDrone, demonstrating its capability for tiny object detection in remote sensing.
Ming-Xue Yang, He Chen, Ning Zhang et al.· IEEE Journal of Selected Top...· 0 citations
Remote sensing object detection suffers from severe performance degradation under cross-domain transfer, where domain gaps arise from differences in spectral response, spatial resolution, and viewing geometry. Most existing unsupervised domain-adaptive object detection (DAOD) methods pursue cross-domain invariance through feature distribution alignment but encounter two limitations specific to remote sensing: cross-domain appearance variation, where differences in imaging conditions produce divergent visual appearances, and foreground-background imbalance, where targets are sparsely distributed across vast backgrounds and alignment is dominated by background statistics. To address these two limitations, frequency-spatial dual-level selective alignment (FS2A), a framework with two complementary modules, is proposed. At the frequency level, scale-class-conditioned frequency modulation (SCFM) decomposes multiscale features via FFT and selectively modulates the low-frequency amplitude conditioned on scale level and categorical composition, restricting adversarial alignment to domain-variant spectral components while preserving the domain-invariant phase spectrum. At the spatial level, kernel relational distillation (KRD) distills pairwise relational structure from a frozen satellite-pretrained vision foundation model (VFM) in polynomial kernel space, where polynomial kernel functions preferentially concentrate alignment on foreground feature pairs over weakly correlated background pairs. Both modules are decoupled from the detection forward pass, introducing no additional inference cost. Experiments on two cross-domain remote sensing benchmarks demonstrate that FS2A achieves 67.6% mAP50 on xView $\rightarrow $ DOTA, surpassing the state-of-the-art by 2.7%, and competitive results on satellite-to-UAV benchmarks. The code will be available at https://github.com/sparklejojo/FSSA-DAOD
Tingting Qiao, He Chen, Jue Wang et al.· IEEE Transactions on Geoscie...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.