Exploring GPT-4o for Semantic Change Detection in Aerial Imagery: An Exploratory Comparison With Traditional and Deep Learning Approaches
Accurate detection of land-use changes from aerial imagery is essential for urban development, environmental monitoring, and infrastructure management. While deep learning has advanced automated change detection, existing solutions remain sensitive to seasonal variations, lighting conditions, and image heterogeneity. This study presents an exploratory workflow-level evaluation of GPT-4o for semantic aerial imagery change detection and compares its behaviour with selected traditional, GIS-assisted, GAN-based, and U-Net-based approaches within the same imagery scenario. Unlike prior work focused on semantic segmentation, we investigate whether a general-purpose multimodal LLM can detect and describe changes without pixel-level training. NDVI exhibited inconsistent class separation; GAN-based mapping achieved a Structural Similarity Index Measure (SSIM) of 0.73 but lacked class fidelity; U-Net produced high accuracy for well-represented classes but struggled to generalize. GPT-4o achieved the best event-level performance, correctly identifying 89.17% of manually annotated changes and providing contextual descriptions and approximate spatial localization. Although promising, LLM performance depends on prompt specification and non-deterministic inference, raising reproducibility challenges. We address these by releasing prompt templates, raw outputs, and controlled inference settings. The results should therefore be interpreted as a case study of one proprietary multimodal model rather than as a comprehensive benchmark of all vision-language models. This exploratory study highlights the emerging potential of multimodal LLMs for interpretable and flexible geospatial analysis while outlining current limitations and future research directions.