Skip to content

GeoSEAN: Explainable Country-Level Image Geolocation for ASEAN Regions

Jul 2026 · arXiv.org · Vol abs/2607.12284 · 0 citations · 28 references
Computer Science

TL;DR

The proposed model can support accurate regional image geolocation while enabling object level inspection of the visual cues underlying its predictions, demonstrating that object frequency and attention based visual evidence capture different aspects of a scene.

Abstract

Image geolocation aims to infer the geographic origin of an image from visual content alone. However, this task remains challenging in regions where countries share similar urban, roadside, architectural, and environmental characteristics. Many existing geolocation models focus on coordinate level prediction or classification performance while providing limited insight into how visual evidence contributes to location predictions. This study presents an explainable country level image geolocation pipeline for 11 ASEAN countries. First, we collected 4,850 images from GeoGuessr style sources, Google Images, and additional street level imagery. We then evaluated three approaches on this dataset: CLIP zero shot classification, a LightGBM classifier, and an MLP classifier. The MLP achieved the best test performance, attaining an accuracy and F1 score of 85.91%. For explainability, predictions generated by the MLP classifier were analyzed post hoc using CLIP attention rollout, YOLO26 object detection on the original images, and Energy Based Pointing Game (EBPG) overlap metrics. Object level analysis indicates that frequently detected objects are not necessarily associated with the highest attention density, suggesting that object frequency and attention based visual evidence capture different aspects of a scene. These results demonstrate that the proposed model can support accurate regional image geolocation while enabling object level inspection of the visual cues underlying its predictions.

View source

Similar papers

Preprint Aug 2026

CoST: Semantic-Aware Urban Understanding via Spatial-Temporal Alignment

Geospatial representation learning from satellite imagery is a fundamental problem for large-scale urban analysis and real-world applications. Despite recent advances, current methods struggle with cross-region generalization and semantic interpretability due to their reliance on region-specific auxiliary data and the neglect of semantic alignment within multi-temporal urban imagery. Therefore, we present CoST, a novel \underline{Co}ntrastive-based \underline{S}patial-\underline{T}emporal framework that aligns spatial context with multi-temporal semantics to extract universal geographic regularities shared across regions. Specifically, CoST explicitly models spatial correlations to capture transferable geographic structures and exploits multi-year urban change semantics to align learned representations with high-level geo-semantics. Extensive experiments demonstrate that CoST consistently achieves superior performance across various downstream tasks and in unseen scenario, yielding an average relative gain of 8.7\% over the strongest competing methods across eight city-indicator settings. The code is available in \href{https://github.com/Arandinglv/CoST}{this repo}.

Yutian Jiang, Jiabo Liu, Xixuan Hao et al. · 1 citation
Conference Jul 2026

Applications of Street View Images based on Artificial Intelligence: A Comprehensive Survey

Street View Images (SVI) are high-resolution, geo-referenced panoramas that capture real-world environments. Integration of Artificial Intelligence (AI) with SVI enables automated analysis for a range of urban applications including object detection, semantic segmentation, text recognition, scene understanding, and socioeconomic prediction. More than 25 recent AI based SVI studies covering applications in crime prediction, building attribute classification, sidewalk inventory, land price estimation, and environmental monitoring were screened for the survey. Across domains, deep learning architectures such as ResNet, ConvNeXt, and Vision Transformers consistently outperformed traditional machine learning models, with reported accuracies up to 94% for classification and R2 values between 0.62–0.83 for prediction tasks. SVI augmented with other data sources like satellite imagery and Global Information System (GIS), enhanced the model performance and contextual understanding. The Findings highlight the dominance of convolutional and transformer-based networks, emerging interest in graph neural networks, and the need for generalized models and diverse datasets to advance SVI research.

Ranjani A, J. C, V. V et al. · 0 citations
Open access Jul 2026

Rapid Georeferencing of Sensor-Limited Helicopter Imagery for Wildfire Response

Abstract. In the initial response to wildfires, securing rapid and accurate geographic information is essential. However, helicopter imagery acquired on-site often lacks precise sensor metadata, such as camera pose and internal parameters, making the application of georeferencing difficult. In particular, obliquely captured wildfire imagery presents additional registration challenges due to severe viewpoint changes, scale variations, and low-texture environments. This study proposes an automated georeferencing pipeline capable of operating under these constraints. The proposed method consists of five stages: preprocessing, image retrieval, feature extraction and matching, Exterior Orientation Parameters (EOP) estimation, and orthomosaic generation. An initial Area of Interest (AOI) is defined using inaccurate initial position data, and the Region of Interest (ROI) within the reference map is obtained through a ResNet50-based image retrieval approach. Subsequently, virtual Ground Control Points (GCPs) are generated through deep learning-based feature matching. Elevation data is then assigned using a Digital Elevation Model (DEM), and EOP are estimated via Perspective-n-Point (PnP) and RANSAC algorithms. Intermediate frames are initialized via interpolation and refined through bundle adjustment to produce the final orthomosaic. Experimental results demonstrated that utilizing SuperGlue and LightGlue complementarily increased the number of successfully georeferenced intervals from 5 to 9. Furthermore, a minimum RMSE of 28.30 m was achieved in the most accurate interval. This method proves that by automating the feature-based georeferencing process, practical geographic information can be rapidly provided for initial disaster response, even in sensor-limited environments.

Seongyun Kim, Jeonghyo Oh, J. Cheon et al. · 0 citations
Open access Aug 2026

GANCIU—Geospatial Analysis with Neural Classification and Image Understanding

GANCIU is introduced, an original hybrid pipeline for the automatic extraction of man-made infrastructure from high-resolution satellite imagery that runs end-to-end on a modest, GPU-free consumer laptop, demonstrating that competitive infrastructure-extraction performance does not require specialised computing hardware.

A. Ganciu, Giovanna Ricci, Margherita Solci · 0 citations
Preprint Aug 2026

What Does CLIP Learn for Regional Geolocalization? Probing Visual Cues and Scene Configuration After Adaptation

Large collections of street-view imagery provide rich visual information about urban environments, but extracting fine-grained geographic information from such data remains challenging. In particular, fine-grained regional geolocalization is challenging because nearby areas often share coarse geographic cues. We study regional geolocalization within a metropolitan area and ask whether pretrained CLIP features are sufficient for regional discrimination, and what visual information supports performance after adaptation. Using 9,085 street-view images from eight Greater Los Angeles regions, we compare zero-shot CLIP, frozen-encoder readouts, partial encoder updating, Low-Rank Adaptation (LoRA), and full fine-tuning. Frozen readouts remain near the 39.03% zero-shot accuracy, whereas encoder adaptation achieves 75.94-82.10%. Full fine-tuning also reduces the mean distance to the predicted region center from 12.30 km to 3.86 km. We probe these gains through semantic cue removal, appearance reduction using edge maps and blur, and scene-configuration disruption using patch scrambling. Adapted models achieve higher edge and blur accuracy and switch 42.92-45.56% of predictions after scrambling, compared with 10.79-14.60% for frozen methods. However, adaptation does not improve the fraction of performance retained after appearance reduction, while vegetation and sky remain influential. A Caltech101 control further shows that scrambling sensitivity is not unique to geolocalization. Overall, encoder adaptation substantially improves nearby-region discrimination and is associated with greater sensitivity to intact scene configuration, without evidence that coarse structure alone becomes sufficient for prediction. These conclusions concern viewpoint variation near known locations rather than geographically disjoint generalization.

Chang-Le Lee, Yeonsoo Park, Abdullah Alfarrarjeh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.