Skip to content
Conference

Research on Intelligent Building Extraction Models Based on High-Resolution Remote Sensing Imagery

Jul 2026 · GEOINFORMATICS · pp. 1-6 · 0 citations · 11 references

Abstract

Building extraction from high-resolution remote sensing imagery is critical for urban planning and smart city development, yet it faces challenges such as blurred boundaries, missing fine details, and severe background interference. To address these issues, this study proposes an improved model named GSU-HRNet, which integrates attention mechanisms and boundary refinement strategies on the basis of UHRNet's high-resolution parallel backbone. An enhanced Pyramid Squeeze-and-Excitation (PSE) module is embedded in the lateral feature transmission paths of each hierarchical stage, capturing multi-scale contextual information via adaptive average pooling of multiple sizes to strengthen semantic responses for buildings and suppress background noise. A Gated Bottleneck Convolution (GBC) module is further introduced in the feature fusion stage, adopting a dual-branch structure with gating mechanisms and residual connections to selectively regulate fused features, alleviate redundant feature accumulation, and improve the stability of feature representation. Experiments were conducted on the aerial imagery subset of the WHU Building Dataset (covering 450 km2 in Christchurch with 8,189 512×512 image tiles), which was split into training, validation and test sets at a ratio of 6:1:3. Ablation experiments verify the effectiveness and complementarity of PSE and GBC modules, with the combined model achieving optimal performance. Quantitative comparisons show that GSU-HRNet outperforms classical models like U-Net and PSPNet, reaching an IoU of 89.93% and an F1-score of 94.50%. Qualitative analysis demonstrates that the proposed model yields clearer building boundaries, more complete structural preservation, and reduced false detections and omissions, even in challenging scenarios with complex building structures and shadow interference. The results confirm that GS-UHRNet effectively enhances feature representation and boundary delineation accuracy, and exhibits strong generalization ability across different building extraction datasets, providing a robust solution for automated building extraction from hig-hresolution remote sensing imagery.

View source

Similar papers

Open access Jul 2026

A Hybrid Pyramid and Strip Pooling Network for Accurate Building Extraction from Remote Sensing Images

SRB-Net is presented, a U-Net-based framework that combines three complementary components: strip pooling for long-range horizontal and vertical context; residual multi-scale atrous spatial pyramid pooling with squeeze-and-excitation blocks for multi-scale and channel-aware feature learning; and a bottleneck attention module (BAM) for refining skip-connection features.

Hamdoun Youssef, Xingyuan Li, Yongtao Yu et al. · 0 citations
Open access Jul 2026

Road Segmentation from Satellite Imagery Based on an Improved SAM Model

Abstract. Road network is an important infrastructure of urban spatial structure and traffic system. Its accurate acquisition is of great significance for urban traffic analysis, automatic driving map construction and disaster emergency response. With the wide acquisition of high-resolution remote sensing images, automatic extraction of road masks from remote sensing images has become an important research direction in the field of remote sensing image understanding. However, the existing deep learning methods still face the problems of obvious modal differences and insufficient modeling of road structure continuity in remote sensing scenes. To solve the above problems, this paper proposes a remote sensing image road segmentation model LR-SAM based on SAM (Segment Anything Model). In this model, the LoRA (Low Rank Adaptation) fine-tuning strategy is introduced to achieve efficient parameter updating, and the MS multi-scale feature interaction module is designed in the coding phase to enhance the expression ability of the linear structure and fine-grained information of the road. At the same time, the original prompt encoder is removed and a lightweight AD decoder is constructed to achieve multi-scale feature fusion. In the reasoning stage, TTA (Test Time Augmentation) strategy is introduced to improve the stability and segmentation accuracy of the model. Experimental results based on CHN6-CUG and SATMTB datasets show that the proposed method achieves 97.20% and 85.06% mIoU and 96.67% and 84.94% F1-score, respectively, which is significantly better than the mainstream road segmentation methods, and verifies the effectiveness of the proposed improvement points.

Bingquan Yao, Haolin Liu, Wenjuan Mao et al. · 0 citations
Open access Aug 2026

ASAR-Net: A Novel Adaptive Scale-Aware Road Extraction Network for High-Resolution Remote Sensing Images

An adaptive scale-aware road extraction network, termed ASAR-Net, which jointly improves multi-scale feature representation and structural continuity and effectively improves both the semantic completeness and structural continuity of extracted road networks is proposed.

Xiaotong Guo, Guang Yang, Yue-bao Wang et al. · 0 citations
Open access Aug 2026

Semantic Segmentation of Remote Sensing Images Based on RS3mamba and Wavelet Transform

The semantic interpretation of remote sensing imagery through segmentation has become indispensable for a wide range of applications, including resource exploration, environmental assessment, and land-use analysis. Yet, accurate parsing of such images remains challenging because complex object boundaries and large scale differences often weaken the ability of conventional Convolutional Neural Network (CNN)-based methods to preserve local details. In response, this study constructs a segmentation framework that couples wavelet convolution with the Mamba architecture. To strengthen feature learning in the intermediate stages, an Auxiliary Segmentation Module (ASM) is employed to provide additional supervisory guidance, which supports optimization and encourages the representation of subtle semantic details. Wavelet-transform convolution is also introduced into the downsampling path, enabling spatial cues and frequency-related information to be exploited in a more coordinated manner for finer boundary and texture modeling. Experiments on public remote sensing datasets and mining area imagery further confirm the effectiveness of the method. Compared with several existing segmentation approaches, the proposed model delivers better overall performance in mIoU, F1-score, and recognition accuracy, particularly in scenes where multiple land-cover categories are heavily interlaced. Moreover, these gains are obtained with relatively low model complexity, suggesting good potential for practical deployment in land monitoring and ecological management.

Wenxi He, Zongmin Yin, Yulong Yang et al. · 0 citations
Open access Sep 2026

Forest Road Extraction from High-Resolution Remote Sensing Imagery Based on an Improved U-Net Model

To address the challenges of vegetation interference, background confusion, and road fragmentation caused by the narrow and elongated structures of forest roads in complex remote sensing imagery, this study proposed an improved U-Net-based model, namely HAA-UNet, for automatic forest road extraction. The proposed model integrates a VGG16 encoder, an Atrous Spatial Pyramid Pooling (ASPP) module, and a Hybrid Dilated Convolution (HDCC) module to enhance feature representation and spatial detail reconstruction of road targets under complex forest environments. Experimental results demonstrated that HAA-UNet achieved superior segmentation performance on the forest road dataset of Xichang City, Sichuan Province, with Precision, Recall, F1-score, and mIoU values of 87.24%, 87.83%, 87.53%, and 79.85%, respectively, outperforming all comparison models. These results indicate that the proposed method effectively improves road continuity and boundary delineation in complex forest scenes. Furthermore, the extracted road information was integrated into the Forest Fire Risk Index (FFRI) assessment framework, demonstrating that accurate road data can improve the spatial characterization of fire risk and provide reliable data support for forest fire risk assessment and forest resource management.

Unknown authors · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.