Skip to content
Open access

A Hybrid Pyramid and Strip Pooling Network for Accurate Building Extraction from Remote Sensing Images

Jul 2026 · Journal of Computing and Electronic Information Management · 0 citations · 48 references

TL;DR

SRB-Net is presented, a U-Net-based framework that combines three complementary components: strip pooling for long-range horizontal and vertical context; residual multi-scale atrous spatial pyramid pooling with squeeze-and-excitation blocks for multi-scale and channel-aware feature learning; and a bottleneck attention module (BAM) for refining skip-connection features.

Abstract

Accurate extraction of building footprints from remote sensing imagery is important for urban planning, disaster management, and geographic information systems. However, complex building shapes, occlusions, and scale variation continue to challenge conventional segmentation models. This paper presents SRB-Net, a U-Net-based framework that combines three complementary components: (1) strip pooling (SP) for long-range horizontal and vertical context; (2) residual multi-scale atrous spatial pyramid pooling (RMASPP) with squeeze-and-excitation (SE) blocks for multi-scale and channel-aware feature learning; and (3) a bottleneck attention module (BAM) for refining skip-connection features. The model was trained with the Adam optimizer and evaluated on the aerial and Satellite Dataset II subsets of the WHU Building Dataset. Among the evaluated baselines, SRB-Net achieved the best overall performance, reaching 98.83% accuracy and 90.12% Intersection over Union (IoU) on the aerial dataset and 98.28% accuracy and 70.89% IoU on Satellite Dataset II. These results show consistent performance improvements across the two evaluated WHU subsets while avoiding claims beyond the within-dataset experimental setting.

Read PDF

Similar papers

Conference Jul 2026

Research on Intelligent Building Extraction Models Based on High-Resolution Remote Sensing Imagery

Building extraction from high-resolution remote sensing imagery is critical for urban planning and smart city development, yet it faces challenges such as blurred boundaries, missing fine details, and severe background interference. To address these issues, this study proposes an improved model named GSU-HRNet, which integrates attention mechanisms and boundary refinement strategies on the basis of UHRNet's high-resolution parallel backbone. An enhanced Pyramid Squeeze-and-Excitation (PSE) module is embedded in the lateral feature transmission paths of each hierarchical stage, capturing multi-scale contextual information via adaptive average pooling of multiple sizes to strengthen semantic responses for buildings and suppress background noise. A Gated Bottleneck Convolution (GBC) module is further introduced in the feature fusion stage, adopting a dual-branch structure with gating mechanisms and residual connections to selectively regulate fused features, alleviate redundant feature accumulation, and improve the stability of feature representation. Experiments were conducted on the aerial imagery subset of the WHU Building Dataset (covering 450 km2 in Christchurch with 8,189 512×512 image tiles), which was split into training, validation and test sets at a ratio of 6:1:3. Ablation experiments verify the effectiveness and complementarity of PSE and GBC modules, with the combined model achieving optimal performance. Quantitative comparisons show that GSU-HRNet outperforms classical models like U-Net and PSPNet, reaching an IoU of 89.93% and an F1-score of 94.50%. Qualitative analysis demonstrates that the proposed model yields clearer building boundaries, more complete structural preservation, and reduced false detections and omissions, even in challenging scenarios with complex building structures and shadow interference. The results confirm that GS-UHRNet effectively enhances feature representation and boundary delineation accuracy, and exhibits strong generalization ability across different building extraction datasets, providing a robust solution for automated building extraction from hig-hresolution remote sensing imagery.

Shi He, Shiye Zhang, Xiujuan Liang et al. · 0 citations
Conference Jul 2026

Multi-Model Evaluation of Semantic Segmentation Techniques for Building Footprint Extraction

In the present generation of increasing geospatial data, accurate and automated extraction of building footprints from high-resolution aerial and satellite imagery has become crucial for various applications such as urban planning, infrastructure development, disaster management, and GIS database maintenance, as manual tracing is time-consuming and unstable for large-scale mapping. This study compares conventional image processing techniques such as thresholding, edge detection, morphological operations through a machine learning approach using Random Forest (RF), and deep learning-based semantic segmentation models, namely U-Net and DeepLabV3+, along with the Segment Anything Model (SAM) using a pre-trained prompt-based setup. All methods are tested on the same set of data, and a standardized data preprocessing is performed for fair comparison. The overall results indicate that the application of DeepLabV3+ is best, with an IoU of 82% and an F1 score of 90%. U-Net achieves second high IoU and F1 scores of 74% and 84% respectively, while Random Forest shows a high IoU of 60% and an F1-score of 72%. SAM has the lowest scores with an IoU of 50% and an F1 score of 51%.

Pravallika Dasapalli, Satya Sahithi, Likitha Kuppila · 0 citations
Open access Sep 2026

Forest Road Extraction from High-Resolution Remote Sensing Imagery Based on an Improved U-Net Model

To address the challenges of vegetation interference, background confusion, and road fragmentation caused by the narrow and elongated structures of forest roads in complex remote sensing imagery, this study proposed an improved U-Net-based model, namely HAA-UNet, for automatic forest road extraction. The proposed model integrates a VGG16 encoder, an Atrous Spatial Pyramid Pooling (ASPP) module, and a Hybrid Dilated Convolution (HDCC) module to enhance feature representation and spatial detail reconstruction of road targets under complex forest environments. Experimental results demonstrated that HAA-UNet achieved superior segmentation performance on the forest road dataset of Xichang City, Sichuan Province, with Precision, Recall, F1-score, and mIoU values of 87.24%, 87.83%, 87.53%, and 79.85%, respectively, outperforming all comparison models. These results indicate that the proposed method effectively improves road continuity and boundary delineation in complex forest scenes. Furthermore, the extracted road information was integrated into the Forest Fire Risk Index (FFRI) assessment framework, demonstrating that accurate road data can improve the spatial characterization of fire risk and provide reliable data support for forest fire risk assessment and forest resource management.

Unknown authors · 0 citations
Open access 2026

Stratified Evaluation of SAM 2 for Zero-Shot Building Segmentation in Aerial Imagery

The first systematic zero-shot evaluation of SAM 2 for aerial building segmentation is presented, establishing SAM 2 as a viable tool for rapid building mapping while highlighting where domain adaptation remains necessary.

Bingning Xiong, Mingyu Ou · 0 citations
Open access Jul 2026

Hierarchical vision mamba U-net for farmland semantic segmentation from remote sensing imagery

Farmland semantic segmentation (FSS) from remote sensing imagery (RSIs) is a critical yet challenging task in precision agriculture. CNN-based methods suffer from limited receptive fields and poor long-range dependency modeling, while Transformers are limited by quadratic computational complexity. To address these issues, this paper proposes a hierarchical Vision Mamba U-Net (HVM-UNet) integrating three key components: a cross-scanning Visual State Space (CSVSS) block to improve scanning performance, a lightweight global feature fusion (GFF) module to replace traditional skip connections for enhanced detail preservation, and a multiscale spatial attention module (MSSA) for refined feature aggregation. Extensive experiments on benchmark datasets demonstrate that HVM-UNet achieves superior segmentation accuracy with linear complexity, outperforming both CNN-based and Transformer-based approaches and offering a robust solution for precision agriculture.

Jing Zhang, Ting Zhang, Guohong Qi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.