2026· IEEE Transactions on Geoscience and Remote Sensing· Vol 64, pp. 5633016-5633016· 0 citations· 70 references
Abstract
Object detection in remote sensing imagery faces challenges such as extreme scale variations and complex backgrounds. Although current methods have made significant strides in visual feature extraction, their predominant focus remains on the image itself, overlooking the potential of integrating external knowledge. To address this limitation, we introduce the knowledge-aware network with region-adaptive fusion for detection (KARFDet), which seamlessly integrates region-specific semantic information with visual features. First, a multiscale fused kernel attention (MSFKA) module is introduced, leveraging a parallel multibranch architecture to enhance contextual feature extraction. Second, a knowledge graph semantic extraction (KGSE) module is designed, employing the random walk with restart (RWR) algorithm to transform discrete knowledge into computable semantic associations. Finally, a novel triple-order knowledge integration (TOKI) mechanism is proposed, which adaptively fuses original, second-order, and probabilistic semantic knowledge, dynamically allocating knowledge weights based on target scale characteristics. Experiments on the DIOR, NWPU VHR-10, and SIMD datasets show that KARFDet achieves mAP50 scores of 65.6%, 91.9%, and 75.8%, respectively, significantly outperforming the baseline model and establishing a new paradigm for semantic-aware detection in complex scenarios. The code is available at https://github.com/ChengXCode/KARFDet
Object detection in remote sensing images is a crucial but challenging research issue in computer vision. Compared to high-resolution images, low-resolution images of the same size typically cover a wider area and thus facilitate efficient object detection. However, the limited visual information and difficulty in distinguishing objects from the background make accurate object detection and localization more challenging. Current object detection networks struggle to adequately extract feature layer information from the feature pyramid in remote sensing images, resulting in unsatisfactory detection performance. To overcome these challenges, we propose a hyper look-ahead network, which incorporates a look-ahead structure (LS), conspicuous feature supplement attention (CFSA), and multiscale feature information process module (MFIPM). The intuition is that the CFSA is integrated into the backbone, enabling the network to rapidly locate objects and effectively ignore continuous blank areas. In addition, by incorporating the LS and the MFIPM in the neck, we enhance the information extraction from different scale feature layers and supplement small object information in deeper features. Experiments on Levir-Ship and VisDrone demonstrate the effectiveness and efficiency of the proposed method. Our method achieves 79.1 mAP with 159 FPS on Levir-Ship, which outperforms many state-of-the-art object detection methods. The code is available online.
Abstract. With the rapid advancement of Urban Air Mobility (UAM), vision-only UAV object detectors like YOLO often suffer from "context blindness" in complex urban canyons, leading to logical fallacies or missed occluded targets. To address these limitations, this paper proposes an innovative Geo-Visual Fusion (GVF) enhancement strategy. By leveraging high-definition (HD) city maps as deterministic geo-spatial priors, we introduce a Geo-spatial Contextual Reasoning (GCR) module to post-process raw visual outputs. This framework incorporates a Semantic Compatibility Matrix (SCM) to eliminate geographically implausible false positives and a Bayesian enhancement rule to boost the confidence of occluded targets. Experimental validation in the Baibuting Community, Wuhan, demonstrates that the GVF framework significantly outperforms the baseline YOLOv11, achieving perfect recall and precision in the test sequence. Furthermore, the 2D vector-based indexing ensures high computational efficiency for edge computing deployment on platforms like the DJI Dock 3. Finally, a closed-loop "reverse empowerment" mechanism for HD map updates is discussed. This work effectively bridges probabilistic computer vision and deterministic geospatial constraints for reliable UAV perception.
Shenman Zhang, Qing-Shan Peng, Meng-Meng Duan et al.· The International Archives o...· 0 citations
A new real-time detector for aerial imagery based on YOLO11, named RSS-YOLO, designed to replace the C3k2 module in YOLO11 and alleviates insufficient integration of spatial and semantic information within the feature extraction layers.
Peng-Fei Dai, Liang Chen, Le Xie et al.· Multimedia Systems· 2 citations
Small object detection in industrial scenarios faces challenges including limited pixel coverage, weak feature representation, and background interference. To address these problems, this paper presents an improved YOLOv11 detection model. First, a dual-backbone network architecture is designed to simultaneously capture rich semantic information and spatial details through parallel feature extraction paths. Second, the SimAM parameter-free attention mechanism is integrated into top-level feature fusion to adaptively enhance features relevant to small objects. Finally, the Adaptive Spatial Feature Fusion (ASFF) module is improved with a dual attention mechanism to optimize multi-scale feature fusion and mitigate feature conflicts. On a self-constructed industrial tool dataset, the method achieves an mAP@0.5:0.95 of 0.920, improving upon the baseline YOLOv11n by 5.9 percentage points. For small object detection specifically, mAP_s reaches 0.898, representing a 7.9 percentage point improvement. Experiments on the public VisDrone dataset further validate the generalization capability of the approach. Results demonstrate that the proposed method significantly enhances small object detection performance, providing an effective solution for industrial vision applications.
Chengru Liu, Junqing Yang, Qi-Qi Guo et al.· 2026 8th International Confe...· 0 citations
An Adaptive and Scalable YOLO model named AS-YOLOR (Adaptive and Scalable YOLO for Rotated object detection), based on the YOLOv8 baseline is proposed, providing a solution with strong practical potential for achieving efficient and high-precision detection of small, rotated objects.
Jin Huang, Juntao Shen, Min Wang et al.· Applied Sciences· 0 citations
An adaptive scale-aware road extraction network, termed ASAR-Net, which jointly improves multi-scale feature representation and structural continuity and effectively improves both the semantic completeness and structural continuity of extracted road networks is proposed.
Xiaotong Guo, Guang Yang, Yue-bao Wang et al.· Applied Sciences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.