Author

Junming Chen

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Open access 2026

A Training-Free Vision Foundation Model Synergy Framework for Building Extraction in High-Resolution Remote Sensing Images

Building extraction holds significant practical importance for urban planning and various human productive activities. However, the complex imaging mechanisms of remote sensing images (RSIs) and the inherent diversity of building features present substantial challenges to accurate extraction. Current mainstream research primarily relies on supervised learning or fine-tuning of foundation models. These approaches depend heavily on large volumes of meticulously annotated data, and thus suffer from high annotation costs and limited generalization capabilities. To overcome these limitations, an unsupervised, training-free framework exploiting pretrained visual foundation models is proposed for building extraction from high-resolution RSIs. This framework requires no human-annotated data, training, or fine-tuning, enabling direct zero-shot inference on high-resolution RSIs. Specifically, it first generates initial pseudolabels by adaptively fusing semantic features from DINO and dense prediction features from CLIP via a spatial correlation-guided weighting mechanism. Then, a dual-path optimization mechanism based on the segment anything model (SAM) is introduced to refine these pseudolabels: local refinement corrects building boundaries, while global filtering enhances regional integrity. This dual-path design, tailored to the geometric characteristics of buildings, is the first attempt to fully exploit SAM’s complementary capabilities in a unified training-free pipeline. Experiments on three public datasets (WHU, WHU-Mix, and Inria) demonstrate that the proposed method achieves F1 scores of 71.86%, 61.48%, and 57.90%, respectively. This performance significantly surpasses current state-of-the-art unsupervised methods, exhibiting excellent accuracy, robustness, and cross-dataset adaptability, and providing valuable insights for practical remote sensing applications.

Junming Chen, Bing Liu, Weiqi Lian et al. · 0 citations