Skip to content
Open access

OSM-CLIP: Enhancing Remote Sensing Image–Text Representation Learning with OpenStreetMap Data

Jul 2026 · Applied Sciences · 0 citations · 16 references

Abstract

Remote sensing vision–language models, such as RemoteCLIP and GeoRSCLIP, have advanced image–text representation learning. However, they rely on manually curated caption datasets that are expensive to scale and provide only global image-level supervision. In this paper, we introduce OSM-CLIP, a framework that exploits the freely available, continuously growing annotations of OpenStreetMap (OSM) to provide regionally scalable, patch-level supervision for remote sensing image-text learning. We construct a large-scale dataset of over 265,000 satellite images covering the contiguous United States, each automatically paired with fine-grained geographic annotations scraped from OSM and mapped to individual image patches. A contrastive loss operating at the patch level associates each image region with its corresponding OSM textual description, enabling the model to learn spatially grounded representations without any manual labeling effort. After fine-tuning on standard remote sensing captioning datasets, OSM-CLIP achieves an average improvement of 10.81% in zero-shot classification, 5.06% in text-to-image retrieval (R@1), and 3.87% in image-to-text retrieval (R@1) over existing methods across 13 classification and 4 retrieval benchmarks. Our results demonstrate that freely available geographic annotations can serve as a powerful source of supervision for remote sensing vision–language models in regions with high-quality OSM coverage.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.