Skip to content
Open access

GANCIU—Geospatial Analysis with Neural Classification and Image Understanding

Aug 2026 · Journal of Imaging · Vol 12 · 0 citations · 69 references
Medicine

TL;DR

GANCIU is introduced, an original hybrid pipeline for the automatic extraction of man-made infrastructure from high-resolution satellite imagery that runs end-to-end on a modest, GPU-free consumer laptop, demonstrating that competitive infrastructure-extraction performance does not require specialised computing hardware.

Abstract

Accurate and up-to-date knowledge of land use and land cover represents one of the central challenges in spatial planning and landscape sciences. In this context, the present work introduces GANCIU (Geospatial Analysis with Neural Classification and Image Understanding), an original hybrid pipeline for the automatic extraction of man-made infrastructure from high-resolution satellite imagery. The primary methodological contribution lies in the sequential integration of four technologically heterogeneous components: a per-pixel Random Forest classifier, a guided image modulation step, edge detection via the Mumford–Shah variational functional solved through the Ambrosio–Tortorelli approximation, and final object delineation via the Segment Anything Model (SAM). Each component does not operate independently but conditions and informs the next: The RF probability map guides the modulation, which in turn directs the sensitivity of the variational step exclusively towards regions of interest; the AT edges provide spatial prompts to SAM, for which its masks are finally filtered by the RF probability in an adaptive manner through a Gaussian Mixture Model. This progressive conditioning scheme constitutes the architectural core of GANCIU and distinguishes it from approaches that combine classification and segmentation in parallel or in purely sequential fashion with each stage conditioning the next but without any reverse correction between them. The Random Forest classifier was trained on 44 manually annotated scenes, geographically disjoint from the twelve independent scenes used for quantitative validation. This validation, based on an instance matching protocol (precision, recall, F1 score, and IoU), confirms the contribution of the full pipeline over a Random-Forest-only baseline: Pooled false positives fall by close to two orders of magnitude (from 8320 to 209), while true positives rise nearly twentyfold (from 5 to 95), with a mean IoU of 0.742 ± 0.060 on correctly matched objects. Notably, the entire pipeline—including SAM-based segmentation—runs end-to-end on a modest, GPU-free consumer laptop (four logical CPU cores, under 16 GB RAM), demonstrating that competitive infrastructure-extraction performance does not require specialised computing hardware.

Read PDF

Similar papers

Open access Jul 2026

Land Cover Classification of Multi-Source Airborne Data using Conventional and Deep-Learning-Based Unsupervised Domain Adaptation

Abstract. For an increasing number of applications, land cover maps can be generated from remote sensing imagery using conventional and deep-learning-based semantic segmentation models. Relying on a large pool of training data, the networks struggle with the spatial-temporal-spectral heterogeneity in the complex and diverse remote sensing imageries, leading to a significant number of errors in the model predictions. This paper presents a workflow comprising domain adaptation and classification. In particular, we analyze two domain adaptation techniques: First, a conventional histogram-matching method, which has turned out to be a surprisingly fast and reliable tool in a previous study, and second, a CycleGAN, which we applied both in its standard form and with the perceptual loss, thereby penalizing style inconsistencies on deeper layers. By applying the workflow to three remote sensing datasets and six directions of domain adaptation, we show that there is “no free lunch” in the sense that all domain adaptation methods have their advantages. Depending on the dataset, classification method, and especially on the availability of 3D data, the performance gap can be reduced to up to 1.5% of the mean F1 score, demonstrating the soundness of the proposed method.

Edwin Deisling, Raphael Zipperer, B. Kottler et al. · 0 citations
Conference Jul 2026

Multi-Model Evaluation of Semantic Segmentation Techniques for Building Footprint Extraction

In the present generation of increasing geospatial data, accurate and automated extraction of building footprints from high-resolution aerial and satellite imagery has become crucial for various applications such as urban planning, infrastructure development, disaster management, and GIS database maintenance, as manual tracing is time-consuming and unstable for large-scale mapping. This study compares conventional image processing techniques such as thresholding, edge detection, morphological operations through a machine learning approach using Random Forest (RF), and deep learning-based semantic segmentation models, namely U-Net and DeepLabV3+, along with the Segment Anything Model (SAM) using a pre-trained prompt-based setup. All methods are tested on the same set of data, and a standardized data preprocessing is performed for fair comparison. The overall results indicate that the application of DeepLabV3+ is best, with an IoU of 82% and an F1 score of 90%. U-Net achieves second high IoU and F1 scores of 74% and 84% respectively, while Random Forest shows a high IoU of 60% and an F1-score of 72%. SAM has the lowest scores with an IoU of 50% and an F1 score of 51%.

Pravallika Dasapalli, Satya Sahithi, Likitha Kuppila · 0 citations
Open access Aug 2026

Semantic Segmentation of Remote Sensing Images Based on RS3mamba and Wavelet Transform

The semantic interpretation of remote sensing imagery through segmentation has become indispensable for a wide range of applications, including resource exploration, environmental assessment, and land-use analysis. Yet, accurate parsing of such images remains challenging because complex object boundaries and large scale differences often weaken the ability of conventional Convolutional Neural Network (CNN)-based methods to preserve local details. In response, this study constructs a segmentation framework that couples wavelet convolution with the Mamba architecture. To strengthen feature learning in the intermediate stages, an Auxiliary Segmentation Module (ASM) is employed to provide additional supervisory guidance, which supports optimization and encourages the representation of subtle semantic details. Wavelet-transform convolution is also introduced into the downsampling path, enabling spatial cues and frequency-related information to be exploited in a more coordinated manner for finer boundary and texture modeling. Experiments on public remote sensing datasets and mining area imagery further confirm the effectiveness of the method. Compared with several existing segmentation approaches, the proposed model delivers better overall performance in mIoU, F1-score, and recognition accuracy, particularly in scenes where multiple land-cover categories are heavily interlaced. Moreover, these gains are obtained with relatively low model complexity, suggesting good potential for practical deployment in land monitoring and ecological management.

Wenxi He, Zongmin Yin, Yulong Yang et al. · 0 citations
Open access Jul 2026

Rapid Georeferencing of Sensor-Limited Helicopter Imagery for Wildfire Response

Abstract. In the initial response to wildfires, securing rapid and accurate geographic information is essential. However, helicopter imagery acquired on-site often lacks precise sensor metadata, such as camera pose and internal parameters, making the application of georeferencing difficult. In particular, obliquely captured wildfire imagery presents additional registration challenges due to severe viewpoint changes, scale variations, and low-texture environments. This study proposes an automated georeferencing pipeline capable of operating under these constraints. The proposed method consists of five stages: preprocessing, image retrieval, feature extraction and matching, Exterior Orientation Parameters (EOP) estimation, and orthomosaic generation. An initial Area of Interest (AOI) is defined using inaccurate initial position data, and the Region of Interest (ROI) within the reference map is obtained through a ResNet50-based image retrieval approach. Subsequently, virtual Ground Control Points (GCPs) are generated through deep learning-based feature matching. Elevation data is then assigned using a Digital Elevation Model (DEM), and EOP are estimated via Perspective-n-Point (PnP) and RANSAC algorithms. Intermediate frames are initialized via interpolation and refined through bundle adjustment to produce the final orthomosaic. Experimental results demonstrated that utilizing SuperGlue and LightGlue complementarily increased the number of successfully georeferenced intervals from 5 to 9. Furthermore, a minimum RMSE of 28.30 m was achieved in the most accurate interval. This method proves that by automating the feature-based georeferencing process, practical geographic information can be rapidly provided for initial disaster response, even in sensor-limited environments.

Seongyun Kim, Jeonghyo Oh, J. Cheon et al. · 0 citations
Jul 2026

GeoSEAN: Explainable Country-Level Image Geolocation for ASEAN Regions

The proposed model can support accurate regional image geolocation while enabling object level inspection of the visual cues underlying its predictions, demonstrating that object frequency and attention based visual evidence capture different aspects of a scene.

Muhamad Syukron, Danish Rafie Ekaputra, T. D. A. Widhianingsih · 0 citations
Review Aug 2026

GeoAI-based post-segmentation quality validation of building footprints via spatial feature engineering

Deep learning-based building footprint extraction from high-resolution imagery often produces topologically inconsistent vectors unfit for direct GIS database ingestion. To address this, we present a multidomain GeoAI quality control framework that automates error detection to systematically purify vector footprint databases. Candidate footprints were generated across five UAV survey sites in Bangladesh using U-Net (ResNet-34) and SAM-LoRA (ViT-B). The extracted raster masks were vectorized, geometrically regularized, and consolidated under a spatial-exclusivity constraint to eliminate duplicate representations. We used twenty-four predictors capturing geometric, spatial-contextual, and raster-derived spectral and texture properties. Machine Learning (ML) classifiers were trained on a development partition (Sites B-D) and rigorously validated on a spatially independent test set (Site E) excluded from hyperparameter tuning and class balancing. The experimental results demonstrate that geometric and spatial-contextual predictors using Decision Tree (DT) provide the most effective discriminatory evidence for identifying object-level boundary deformations. DT achieved an accuracy of 95.31%, an F1-score of 91.06%, and a Matthews correlation coefficient (MCC) of 0.880 on the unseen testing site. At the database level, this framework successfully identified 87.34% of erroneous footprints while maintaining 98.31% of acceptable structures, reducing the residual error proportion from 27.32% to 4.62% and improving final database purity to 95.38%. This translates into a relative error reduction of 83.09%. The findings indicate that post-segmentation object-level ML provides a highly transferable, robust mechanism for automated quality assurance in production-ready geographic information system (GIS) workflows.

Shah Imran Ahsan Chowdhury, Kazi Jihadur Rashid, Rajsree Das Tuli et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.