Jul 2026· ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences· Vol XI-4-2026, pp. 103-110· 1 citation· 8 references
TL;DR
These findings demonstrate the complementary value of integrating spectral and structural information for robust building footprint extraction and how domain adaptation strategies can be used to enhance cross-regional transferability.
Abstract
Abstract. Accurate building footprint extraction is critical for applications ranging from population estimation to disaster management. Although optical imagery provides detailed spectral information, it often struggles with shadows, occlusions, and background clutter in dense urban environments. Lidar data, by contrast, offer precise elevation and structural attributes but face challenges such as variable point density and noise. This study integrates multispectral imagery from the U.S. Department of Agriculture (USDA) National Agriculture Imagery Program (NAIP) with lidar-derived feature height and intensity from the U.S. Geological Survey (USGS) 3D Elevation Program (3DEP) to improve footprint extraction using a U-Net–based deep learning model. A six-band input stack (RGB, near-infrared, height, intensity) was developed, normalized, and tiled for training and evaluation against Microsoft Global Building Footprints (GBF). Results from the Houston, TX test site show that the six-band model achieved a precision of 0.86, recall of 0.88, F1 score of 0.87, and Intersection-over-Union (IoU) of 0.76, consistently outperforming four-band baselines by reducing false positives while maintaining sensitivity. Predictions on withheld Houston tiles confirmed strong within-region generalization, yielded a precision of 0.78, recall of 0.81, F1 score of 0.79, and IoU of 0.66. Qualitative analysis further revealed limitations stemming from both training label quality and vegetation–building confusion. These findings demonstrate the complementary value of integrating spectral and structural information for robust building footprint extraction and how domain adaptation strategies can be used to enhance cross-regional transferability.
Three-dimensional (3D) building reconstruction plays an important role in urban planning, disaster resilience, and sustainable infrastructure development. Conventional approaches based on airborne Light Detection and Ranging (LiDAR) data can be computationally intensive and often require large labeled datasets, particularly for large-scale applications. This study presents a semi-automated workflow for reconstructing multi-level 3D building models by integrating airborne LiDAR point cloud data with building footprints extracted from National Agriculture Imagery Program (NAIP) imagery using a Mask Region-Based Convolutional Neural Network (Mask R-CNN). The extracted footprints were used to spatially isolate building-specific LiDAR subsets for 3D reconstruction. The proposed methodology generated building models at multiple Levels of Detail (LOD), ranging from two-dimensional (2D) footprints to volumetric representations with detailed roof structures. Building footprint extraction was quantitatively evaluated against LiDAR-derived footprints generated using Density-Based Spatial Clustering of Applications with Noise (DBSCAN), which served as the reference dataset and detected buildings obscured by tree canopy. The reconstruction workflow was implemented in Open3D and incorporated a boundary-aware mesh refinement strategy based on ear-clipping triangulation to improve rooftop continuity in LOD2 models. The proposed footprint extraction framework achieved a mean Intersection over Union (IoU) of 0.8226 relative to LiDAR-derived reference building footprints, indicating reliable building delineation that supports the proposed multi-level 3D building reconstruction workflow.
L. Lakshmanan, S. Nagarajan· Urban Science· 0 citations
SRB-Net is presented, a U-Net-based framework that combines three complementary components: strip pooling for long-range horizontal and vertical context; residual multi-scale atrous spatial pyramid pooling with squeeze-and-excitation blocks for multi-scale and channel-aware feature learning; and a bottleneck attention module (BAM) for refining skip-connection features.
Hamdoun Youssef, Xingyuan Li, Yongtao Yu et al.· Journal of Computing and Ele...· 0 citations
Accurate extraction of the spatial distribution of buildings from remote sensing imagery in complex urban environments is essential for urban planning and development. However, existing methods often suffer from high computational costs and insufficient building boundary recovery, making it difficult to achieve both efficient and accurate building extraction. To address these limitations, this study proposes LGAS-UNet, a lightweight network for building extraction. Based on the UNet architecture, LGAS-UNet replaces the original encoder with LSNet and incorporates a Global Context Aggregation Module (GCAM), Attention Gates (AGs), and Self-Calibrated Convolution (SCConv) modules into the encoder–decoder bridge, skip connections, and decoder feature-fusion units, respectively. These components enhance the global contextual representation of deep features, suppress irrelevant background responses during cross-level feature propagation, and improve feature fusion and boundary detail recovery during decoding. Experiments were conducted on the public WHU Building Dataset and a Zhengzhou building dataset constructed from satellite imagery. With only 6.12 M parameters and 3.64 G FLOPs, LGAS-UNet achieved an intersection over union (IoU) of 86.24%, an F1-score of 92.61%, and a boundary F1-score (BF-score) of 87.84% on the WHU dataset, achieving the best overall performance among the compared methods. On the Zhengzhou building dataset, LGAS-UNet achieved an IoU of 72.37%, an F1-score of 83.97%, and a BF-score of 74.09%, representing improvements of 1.70, 1.16, and 2.11 percentage points, respectively, over UNet. These results demonstrate that LGAS-UNet can efficiently and accurately extract buildings from remote sensing imagery in complex urban environments, providing a practical methodological reference for urban planning and management.
Yao Lu, Gang Cheng, Guo-Sheng Cai et al.· Italian National Conference...· 0 citations
Abstract. Semantic classification is a fundamental step in Mobile Laser Scanning (MLS) point clouds processing, and remains a non-trivial task. In this work, we propose a classification framework based on a 3D Sparse Convolutional Neural Network (SparseCNN) for efficient processing of large-scale MLS data. A coarse-to-fine two-stage pipeline is introduced, where an essential model performs a classification for the entire scene, followed by a refinement stage for detailed ground-surface classes. To enhance robustness under diverse acquisition conditions, both point-wise and scene-wise data augmentation strategies are employed during the training, including rotation, jittering, density perturbation, noise injection, and patch swapping. To account for environmental and sensor variations, wavelength-specific models are trained for both urban and highway scenes. Experimental results on urban and highway datasets demonstrate strong performance, achieving over 90% accuracy for major classes, while ablation studies show that radiometric features are critical for distinguishing material dependent classes, such as traffic signs, and that the proposed augmentation strategies improve performance for challenging object categories, such as pedestrian, which is dynamic and structurally ambiguous.
Nan-Feng Li, H. Teufelsbauer, F. Pöppl et al.· The International Archives o...· 0 citations
Accurate building height information at the individual footprint scale is essential for material stock accounting and post-disaster damage assessments yet remains difficult to obtain at city scale in the Global South where airborne LiDAR coverage is rare and commercial very high-resolution imagery is cost-prohibitive or unavailable. While recent works have demonstrated building height estimation using freely available Sentinel imagery, the resolution ceiling of resulting products is still coarse for material stock analysis. This study incorporates products derived from data freely accessible under scientific research licenses, TerraSAR-X StripMap and PlanetScope, alongside Sentinel-1 to predict building heights in a large city in Brazil. To account for the spatial autocorrelation in the training set, features from all sources are integrated in a geographically weighted random forest model, returning an RMSE of 5.34 m and R2 of 0.756 against a LiDAR reference dataset. Local feature importance showed predictor dominance to vary consistently across intra-urban contexts, with footprint geometry dominating for low-rise buildings, shadow-derived height for taller and more isolated structures, and spectral reflectance for the tallest buildings in the set. Sentinel-1 backscatter and InSAR occupy complementary spatial niches, with no single sensor uniformly preferable across the set. Results provide optioneering guidance and insight over satellite-derived products predictive relevance in distinct contexts, which global machine learning or neural network models cannot offer.
Guilherme Iablonovski, P. Frison, T. D. da Silva· 0 citations
In the present generation of increasing geospatial data, accurate and automated extraction of building footprints from high-resolution aerial and satellite imagery has become crucial for various applications such as urban planning, infrastructure development, disaster management, and GIS database maintenance, as manual tracing is time-consuming and unstable for large-scale mapping. This study compares conventional image processing techniques such as thresholding, edge detection, morphological operations through a machine learning approach using Random Forest (RF), and deep learning-based semantic segmentation models, namely U-Net and DeepLabV3+, along with the Segment Anything Model (SAM) using a pre-trained prompt-based setup. All methods are tested on the same set of data, and a standardized data preprocessing is performed for fair comparison. The overall results indicate that the application of DeepLabV3+ is best, with an IoU of 82% and an F1 score of 90%. U-Net achieves second high IoU and F1 scores of 74% and 84% respectively, while Random Forest shows a high IoU of 60% and an F1-score of 72%. SAM has the lowest scores with an IoU of 50% and an F1 score of 51%.