Skip to content

Author

Mohammed Alruqimi

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

OSM-CLIP: Enhancing Remote Sensing Image–Text Representation Learning with OpenStreetMap Data

Remote sensing vision–language models, such as RemoteCLIP and GeoRSCLIP, have advanced image–text representation learning. However, they rely on manually curated caption datasets that are expensive to scale and provide only global image-level supervision. In this paper, we introduce OSM-CLIP, a framework that exploits the freely available, continuously growing annotations of OpenStreetMap (OSM) to provide regionally scalable, patch-level supervision for remote sensing image-text learning. We construct a large-scale dataset of over 265,000 satellite images covering the contiguous United States, each automatically paired with fine-grained geographic annotations scraped from OSM and mapped to individual image patches. A contrastive loss operating at the patch level associates each image region with its corresponding OSM textual description, enabling the model to learn spatially grounded representations without any manual labeling effort. After fine-tuning on standard remote sensing captioning datasets, OSM-CLIP achieves an average improvement of 10.81% in zero-shot classification, 5.06% in text-to-image retrieval (R@1), and 3.87% in image-to-text retrieval (R@1) over existing methods across 13 classification and 4 retrieval benchmarks. Our results demonstrate that freely available geographic annotations can serve as a powerful source of supervision for remote sensing vision–language models in regions with high-quality OSM coverage.

Alessio Pierdominici, R. Ricci, Mohammed Alruqimi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.