Skip to content
Open access

Building Segmentation from High-Resolution Remote Sensing Imagery: A Comparative Evaluation of Convolutional and Transformer Architectures

Sep 2026 · Gazi Üniversitesi Fen Bilimleri Dergisi Part C: Tasarım ve Teknoloji · 0 citations · 12 references

TL;DR

The results show that model selection in building extraction cannot rest on a single accuracy figure, and that boundary-based metrics should govern the choice in applications where the position of the boundary is critical.

Abstract

Architectures for building extraction are usually compared under different training and inference set-ups, which makes the reported performance figures incommensurable. This study evaluates three convolutional segmentation architectures (U-Net, U-Net++, DeepLabV3+) and one transformer-based architecture (SegFormer) on the aerial orthophoto subset of the WHU Building Dataset under a single training and inference protocol. Fixing the encoder to ResNet-34 in the convolutional models separates the effect of decoder design from backbone capacity. U-Net++ achieved the highest performance at 89.04% IoU and 94.20% F1, and the difference between the models remained within a narrow range of 2.42 points on the standard IoU metric. On the boundary IoU metric, which measures overlap only within the edge band, the same spread rose to 9.26 points, roughly 3.8 times as wide. Reporting standard IoU alone therefore understates the difference between architectures systematically. An inference protocol combining an exponential moving average, eight-way test-time augmentation and a decision threshold tuned on the validation set yielded gains of 0.47 to 0.69 IoU points without changing the ranking of the models; that magnitude corresponds to roughly a quarter of the total spread attributable to architecture. On the cost side, parameter count proved a misleading indicator: a 6.7% difference in parameters between U-Net++ and U-Net corresponded to a 135.2% difference in multiply-accumulate operations. The results show that model selection in building extraction cannot rest on a single accuracy figure, and that boundary-based metrics should govern the choice in applications where the position of the boundary is critical.

Read PDF

Similar papers

Differential morphological profile neural networks for segmentation of remote sensing imagery

Semantic segmentation, assigning a class label to each pixel, has been revolutionized by deep neural networks. A significant milestone, the Fully Convolutional Network (FCN), demonstrated that a purely convolutional architecture could outperform previous approaches. Subsequent architectures largely adopted its encoder–...

David Huangal · 0 citations
Open access Sep 2026

An Improved DeepLabV3+ Network for Bare Rock Segmentation in Remote Sensing Images

To improve the binary semantic segmentation of bare rock and background in complex urban environments, this study developed an improved DeepLabV3+ model for red–green–blue (RGB) imagery derived from Gaofen-1 (GF-1) data. Four backbone networks—Xception, ResNet50, MobileNetV2, and MobileNetV3—were first evaluated. ResNe...

Qiang Wang, Hao-Chuan Lei, Xia-Song Hu et al. · 0 citations
2026

Three-Branch Hybrid Network for Farmland Segmentation in Remote Sensing Images

The accurate segmentation of remote sensing imagery is critical for precision agriculture but challenging due to spectral complexity and ambiguous interclass boundaries. The convolutional neural networks are limited in modeling global context, while transformer-based methods incur high computational overhead. This lett...

Wei-Hui Zeng, Fang Wang, Gensheng Hu · 0 citations
Open access Sep 2026

LGAS-UNet: A Lightweight Network for Building Extraction from Remote Sensing Imagery in Complex Urban Scenes

Accurate extraction of the spatial distribution of buildings from remote sensing imagery in complex urban environments is essential for urban planning and development. However, existing methods often suffer from high computational costs and insufficient building boundary recovery, making it difficult to achieve both ef...

Yao Lu, Gang Cheng, Guo-Sheng Cai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.