Skip to content

Collaborative Context-Affine Perception Network for Remote Sensing Small Object Detection

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 5641014-5641014 · 0 citations · 66 references

Abstract

Small object detection in remote sensing images (RSIs) is challenging because imaging degradation weakens object textures, reduces contrast, and blurs boundaries. These effects are further aggravated by hierarchical feature extraction, where repeated downsampling weakens shallow spatial cues before they reach deeper semantic representations. To alleviate these effects, we propose the collaborative context-affine perception network (CoCAPNet), which progressively refines structural, contextual, and spatial representations along the detection pipeline. CoCAPNet comprises four complementary components. First, the context-aware dynamic affine module (CDAM) separates feature responses into structure- and detail-dominant components and adaptively recombines them to retain weak spatial cues. Next, the dual-range perception fusion module (DPFM) adaptively balances local detail and global semantics by integrating scale-dependent response predictions with multilevel receptive field processing. The collaborative multikernel block (CMKB) then addresses sparse and unstable activations in small objects by enhancing object-relevant patterns via dynamic multikernel perception, fine-grained refinement, and channel-aware weighting. Finally, the fine-grained small object detector (FSDetector) introduces a high-resolution prediction branch to retain spatial detail for small targets. These components form a progressive feature-processing pipeline from representation reconstruction to scale-aware fusion, structural refinement, and high-resolution prediction. Experiments on the DIOR, NWPU VHR-10, and AI-TOD yield mean average precision (mAP) values of 86.3%, 95.8%, and 59.1%, respectively, indicating consistent gains across datasets with substantial scale variation and densely distributed small objects.

View source

Similar papers

Aug 2026

Context-gated dynamic perception for small-object detection in dense aerial scenes

A context-gated dynamic perception framework that treats small-object feature degradation as a coupled problem of representation, fusion, and prediction and indicates a practical accuracy-efficiency trade-off for dense aerial small-object perception.

Guang-Jun Gao, Ruibing Xie · 0 citations
2026

SPEFormer: A Synergistic Perception-Enhanced Transformer With Density-Guided Dynamic Queries for Remote Sensing Small-Object Detection

Small-object detection in remote sensing images is highly challenging due to their limited pixel representation, weak feature expression, and strong interference from complex backgrounds. Moreover, existing query-based detectors typically use a fixed number of object queries for all images, making them poorly adaptive...

Ye Yuan, Shuang-Long Li, Jian Liu et al. · 0 citations
2026

Frequency–Spatial Joint Decoupling With Adaptive Perceptual Aggregation for Remote Sensing Small Object Detection

Small object detection in remote sensing (RS) imagery remains fundamentally challenging due to severe information degradation caused by limited spatial resolution and complex background interference. In deep neural networks, such degradation is further exacerbated by irreversible information loss during conventional do...

Ying Gao, Zongshuai Zhang, Zheng-Yu Zhu et al. · 0 citations
2026

Remote Sensing Object Detection Based on Detail-Semantic Decoupling and Multiscale Coordinate-Guided Semantic Enhancement

Due to the severe scale variation of targets in remote sensing images, the dense distribution of objects, and the fact that many small targets occupy only a very limited number of pixels, existing detection methods are prone to losing shallow details and suffering from insufficient low-level semantic representation dur...

Fa-Quan Song, Wu Le, Ming Lv et al. · 1 citation
Preprint Aug 2026

SPEANet: Structural Prior Enhanced Attention Network for Parameter-Efficient Remote Sensing Object Detection

Remote sensing object detection (RSOD) requires compact backbones capable of preserving weak geometric cues under extreme scale variation and background clutter. Fixed structural operators provide complementary contour and frequency responses without introducing learnable operator coefficients. However, directly inject...

Wei Lu, Jun-Jie Li, Fei-Fei Sang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.