2026· IEEE Transactions on Geoscience and Remote Sensing· Vol 64, pp. 5641014-5641014· 0 citations· 66 references
Abstract
Small object detection in remote sensing images (RSIs) is challenging because imaging degradation weakens object textures, reduces contrast, and blurs boundaries. These effects are further aggravated by hierarchical feature extraction, where repeated downsampling weakens shallow spatial cues before they reach deeper semantic representations. To alleviate these effects, we propose the collaborative context-affine perception network (CoCAPNet), which progressively refines structural, contextual, and spatial representations along the detection pipeline. CoCAPNet comprises four complementary components. First, the context-aware dynamic affine module (CDAM) separates feature responses into structure- and detail-dominant components and adaptively recombines them to retain weak spatial cues. Next, the dual-range perception fusion module (DPFM) adaptively balances local detail and global semantics by integrating scale-dependent response predictions with multilevel receptive field processing. The collaborative multikernel block (CMKB) then addresses sparse and unstable activations in small objects by enhancing object-relevant patterns via dynamic multikernel perception, fine-grained refinement, and channel-aware weighting. Finally, the fine-grained small object detector (FSDetector) introduces a high-resolution prediction branch to retain spatial detail for small targets. These components form a progressive feature-processing pipeline from representation reconstruction to scale-aware fusion, structural refinement, and high-resolution prediction. Experiments on the DIOR, NWPU VHR-10, and AI-TOD yield mean average precision (mAP) values of 86.3%, 95.8%, and 59.1%, respectively, indicating consistent gains across datasets with substantial scale variation and densely distributed small objects.
A context-gated dynamic perception framework that treats small-object feature degradation as a coupled problem of representation, fusion, and prediction and indicates a practical accuracy-efficiency trade-off for dense aerial small-object perception.
Guang-Jun Gao, Ruibing Xie· Pattern Analysis and Applica...· 0 citations
Small-object detection in remote sensing images is highly challenging due to their limited pixel representation, weak feature expression, and strong interference from complex backgrounds. Moreover, existing query-based detectors typically use a fixed number of object queries for all images, making them poorly adaptive...
Ye Yuan, Shuang-Long Li, Jian Liu et al.· IEEE Transactions on Geoscie...· 0 citations
Small object detection in remote sensing (RS) imagery remains fundamentally challenging due to severe information degradation caused by limited spatial resolution and complex background interference. In deep neural networks, such degradation is further exacerbated by irreversible information loss during conventional do...
Ying Gao, Zongshuai Zhang, Zheng-Yu Zhu et al.· IEEE Transactions on Geoscie...· 0 citations
Due to the severe scale variation of targets in remote sensing images, the dense distribution of objects, and the fact that many small targets occupy only a very limited number of pixels, existing detection methods are prone to losing shallow details and suffering from insufficient low-level semantic representation dur...
Fa-Quan Song, Wu Le, Ming Lv et al.· IEEE Transactions on Geoscie...· 1 citation
Remote sensing object detection (RSOD) requires compact backbones capable of preserving weak geometric cues under extreme scale variation and background clutter. Fixed structural operators provide complementary contour and frequency responses without introducing learnable operator coefficients. However, directly inject...
Wei Lu, Jun-Jie Li, Fei-Fei Sang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.