Skip to content
Open access

From Aerial to Satellite: Can Super-Resolution Enable Label-Free Model Transfer?

Jul 2026 · ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences · Vol XI-2-2026, pp. 493-501 · 0 citations · 13 references

TL;DR

This work investigates whether super-resolution (SR) methods can bridge the gap between aerial and high-resolution satellite imagery, enabling a label-free model transfer, meaning without fine-tuning the authors' model with additional manual annotations.

Abstract

Abstract. Satellite imagery enables large-scale remote sensing applications by providing frequent and large-scale coverage. However, its limited spatial resolution often restricts the use of satellite images in tasks that require detailed, fine-scale information. In contrast, aerial images offer a much higher spatial resolution, allowing the extraction of fine-grained features, but typically cover smaller, more localized areas. In this work, we investigate whether super-resolution (SR) methods can bridge the gap between aerial and high-resolution satellite imagery, enabling a label-free model transfer, meaning without fine-tuning our model with additional manual annotations. The idea is to enhance the spatial resolution of high-resolution satellite images, allowing models trained on aerial data to be directly applied to satellite images. Towards this goal, a state-of-the-art SR algorithm is used to upscale three high-resolution satellite images, matching the resolution of the aerial training data. Then, a segmentation network trained on an aerial image dataset is applied to segment roads and parking areas in the super-resolved satellite images. The approach is evaluated on an annotated dataset and compared to the results in the original satellite images. Additionally, we investigate its performance on a low-resolution aerial image. Our results demonstrate that SR facilitates the utilization of models trained on aerial image datasets for large-scale satellite applications without requiring new labels.

Read PDF

Similar papers

#generative ai Sep 2026

Satellite imagery super-resolution using GANs and aerial images

Satellite imagery often suffers from limited spatial resolution and, in many cases, high acquisition costs. These factors restrict their use in applications such as urban monitoring, land management, and wildlife studies. This work proposes an AI-based super-resolution approach that leverages high resolution aerial imagery to train a Generative Adversarial Network. Specifically, the ESRGAN (Enhanced Super-Resolution Generative Adversarial Network) architecture is adapted and trained using aerial orthophotos, enabling the transfer of learned spatial representations to low-resolution satellite images. The trained model is evaluated on satellite image patches at 2 and 4 super-resolution scales. Performance is assessed using structural, perceptual, and chromatic metrics, including SSIMY, MS-SSIM, LPIPS and CIEDE2000. The results show clear improvements, with increased sharpness, enhanced edge definition, and consistent reconstruction of urban structures and terrain features. From a quantitative perspective, the 2 scale achieves the best overall metric values, while the 4 scale maintains stable and meaningful performance despite the higher reconstruction difficulty. These findings demonstrate the feasibility of transferring super-resolution capabilities from aerial images to satellite imagery, even in the presence of spectral and geometric differences between acquisition domains. Overall, this study provides a solid foundation for the development of low-cost, AI-driven satellite image super-resolution models and outlines future research directions focused on dataset expansion, domain adaptation strategies, and sensor-specific architectural improvements.

Magda Alexandra Trujillo-Jiménez, Francisco Iaconis, Debora Pollicelli et al. · 0 citations
Open access Jul 2026

From Super-Resolution to Superior Land Cover Detection: Cross-Channel Attention Network for Aerial Image

MAPSRNet offers a practical solution for scenarios where HR imagery is limited or unavailable, highlighting its potential for large-scale remote sensing applications and demonstrating that perceptual and structural fidelity, rather than pixel-level similarity, can drive superior performance in urban land cover segmentation.

Yuwei Cai, Zhimeng He, Meiliu Wu et al. · 0 citations
Open access Jul 2026

A Dynamically Weighted Framework for Adaptive Reference-Based Super-Resolution

Abstract. Satellite remote sensing is inherently constrained by a trade-off between spatial and temporal resolution. As a result, high-temporal-frequency sensors such as Geostationary Ocean Color Imager-II provide operationally valuable observations but at coarse spatial resolution. Reference-Based Super-Resolution (Ref-SR) can address this limitation by transferring high-resolution textures from an external reference image, but temporal mismatch between the target and reference images often leads to unreliable texture transfer and severe artifacts. This problem becomes more critical in extreme low-resolution (LR) settings, where structural information is already severely degraded. To address this issue, we propose the Dynamic Ref-SR Framework, which computes a pixel-wise weight map from intensity differences between the LR and reference images to selectively control reference transfer. The resulting weights promote reference use in stable regions while suppressing it in temporally inconsistent regions. The framework was validated on three backbone architectures—CNN (EDSR), Swin Transformer, and GAN—using a Sentinel-2 dataset for four-band reconstruction (RGB and NIR). Across all metrics and architectures, the proposed Ref-SR framework consistently outperformed the SISR baseline in both structural and spectral evaluations. Among the tested backbones, the GAN-based model achieved the best overall performance, with a PSNR of 35.60 dB, an SSIM of 0.92, a SAM of 2.20°, and an ERGAS of 74.71. These results demonstrate that the proposed framework can improve LR satellite imagery while reducing the risk of reference misuse under temporal mismatch.

Chae-Eun Kim, Junhwa Chi · 0 citations
Open access Jul 2026

Land Use and Land Cover Classification Using Transfer Learning and Temporal Convolutional Networks on Low-Resolution Remote Sensing Images

Recently, low-resolution remote sensing (RS) images have received significant attention because of their widespread spatial coverage, minimum acquisition cost, quick transmission ability, and large-scale earth observation suitability. However, land-use and land-cover (LULC) classification using low-resolution satellite imagery remains challenging due to restricted spatial information, spectral similarity amongst land-cover classes, noise differences, and complex scene heterogeneity. Though recent deep learning-based models have exhibited effective outcomes, they still suffer from insufficient feature representation, inadequate contextual dependency learning, and minimal classification accuracy when processing low-resolution RS images. To resolve these issues, this study develops a lightweight feature extraction model with Temporal Convolutional Networks for low-resolution remote sensing image classification. The proposed model initially preprocesses the images to improve feature consistency and quality. The feature extraction phase then employs MobileNet-V2 to identify and represent relevant spatial patterns in RS images, followed by a temporal convolutional network for RSI classification, enabling effective modeling of sequential and contextual dependencies in spatial features. Furthermore, adaptive fine-tuning of model parameters is performed using an artificial rabbit optimization algorithm to enhance classification accuracy and convergence behavior. Extensive experimental evaluation of the LFEARO-LULCRSI model on the benchmark EuroSat Dataset from Sentinel-2 imagery demonstrates improved performance over existing methods, achieving an accuracy of 98.57%. An ablation study is also performed to examine the contribution of individual model components. The proposed model thus proves useful for effective geospatial analysis in agriculture, urban planning, disaster assessment, and sustainable environmental management, enhancing feature discrimination and contextual dependency learning in low-resolution satellite imagery.

G. Sravanthi, A. Gnanasekaran, G. Ramesh · 0 citations
Open access Aug 2026

Towards Lightweight and Accurate Remote-Sensing Image Super-Resolution via Reparameterized Feature Enhancement Network

Remote sensing image super-resolution (RSISR) provides an effective means of improving spatial detail for Earth observation and satellite image interpretation. However, existing methods often rely on increasingly complex network designs with deeper hierarchies and expanded channel capacities to pursue higher performance, resulting in heavy models with high computational cost, which restricts their deployment on resource-constrained platforms. To address this challenge, we propose a novel reparameterized feature enhancement network (RepFEN) for lightweight and accurate RSISR tasks. Specifically, a multi-scale reparameterized module (MRepM) is designed to capture multi-scale spatial information and enhance texture representation. Furthermore, a partial-channel gated attention module (PCGAM) is introduced to selectively enhance discriminative features along the channel dimension, effectively improving fine-grained detail restoration. By integrating structural reparameterization and multi-scale lightweight modules, the proposed method achieves a better balance between reconstruction accuracy and inference efficiency. Extensive experiments on both remote sensing and natural image super-resolution benchmarks demonstrate that our method achieves superior performance compared to existing state-of-the-art methods, while maintaining minimal computational overhead, showing significant potential for real-world applications.

Feng Huang, Ren-Hui Wei, Liqiong Chen et al. · 0 citations
Preprint Jul 2026

Self-supervised training for high-resolution close-range multispectral remote sensing imagery

Although self-supervised learning (SSL) offers a promising way to reduce annotation effort in close-range remote sensing, its effectiveness for high-resolution multispectral unmanned aerial vehicle (UAV) imagery remains underexplored due to limited data. This study evaluated SSL pretraining for precision agriculture using cm-scale multispectral drone imagery collected across multiple sensors, years, and regions. Transformer-based encoders were pretrained with Momentum Contrast v3 (MoCo-v3) and Masked Autoencoders on a harmonized dataset combining msuav500K with newly collected multi-year UAV imagery from agricultural fields in Finland. Pretraining used four spectral bands (Green, Red, Red-Edge, Near-Infrared) for cross-sensor compatibility. The models were evaluated on crop-weed semantic segmentation using the WeedMap dataset with 5--100% training data. The following two subsets served as downstream tasks: Task A (Germany, RedEdge-M), where all pretrained models were compared under partial and full fine-tuning, and Task B (Switzerland, Sequoia), where the best encoder from Task A was assessed. Our Swin Transformer pretrained with MoCo-v3 achieved the strongest performance on both tasks, surpassing the Swin Transformer model of Doornbos et al. pretrained on a pre-release of msuav500K. Our pretrained Swin Transformer further demonstrated cross-sensor and cross-region generalization. We additionally provide a public multi-year multispectral UAV dataset from Finland to support future research.

Leon-Friedrich Thomas, Mikael Anakkala, Antti Lajunen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.