LGAS-UNet: A Lightweight Network for Building Extraction from Remote Sensing Imagery in Complex Urban Scenes
Abstract
Accurate extraction of the spatial distribution of buildings from remote sensing imagery in complex urban environments is essential for urban planning and development. However, existing methods often suffer from high computational costs and insufficient building boundary recovery, making it difficult to achieve both efficient and accurate building extraction. To address these limitations, this study proposes LGAS-UNet, a lightweight network for building extraction. Based on the UNet architecture, LGAS-UNet replaces the original encoder with LSNet and incorporates a Global Context Aggregation Module (GCAM), Attention Gates (AGs), and Self-Calibrated Convolution (SCConv) modules into the encoder–decoder bridge, skip connections, and decoder feature-fusion units, respectively. These components enhance the global contextual representation of deep features, suppress irrelevant background responses during cross-level feature propagation, and improve feature fusion and boundary detail recovery during decoding. Experiments were conducted on the public WHU Building Dataset and a Zhengzhou building dataset constructed from satellite imagery. With only 6.12 M parameters and 3.64 G FLOPs, LGAS-UNet achieved an intersection over union (IoU) of 86.24%, an F1-score of 92.61%, and a boundary F1-score (BF-score) of 87.84% on the WHU dataset, achieving the best overall performance among the compared methods. On the Zhengzhou building dataset, LGAS-UNet achieved an IoU of 72.37%, an F1-score of 83.97%, and a BF-score of 74.09%, representing improvements of 1.70, 1.16, and 2.11 percentage points, respectively, over UNet. These results demonstrate that LGAS-UNet can efficiently and accurately extract buildings from remote sensing imagery in complex urban environments, providing a practical methodological reference for urban planning and management.