Skip to content
Open access

Diffusion-Driven Super-Resolution for Remote Sensing: Dual Decoder Attention Models for High-Fidelity Earth Observation

Sep 2026 · International Journal of Engineering and Modern Technology · 0 citations

Abstract

Remote sensing images have played a significant role in various areas in recent years, such as urban infrastructure planning, disaster detection, and environmental detection. They provide helpful information about pollution levels, climate change, and farming conditions which helps to take timely actions that enhance efficiency and reduce waste. Also, farmers can rapidly evaluate estimable actions to enhance efficiency while reducing the wastage of resources. Nevertheless, atmospheric interference, sensor limitations, and spatial resolution gaps can often compromise the quality of these images. Advanced technologies like deep learning techniques can, fortunately, solve this problem. This study applies a diffusion model for denoising and uses a U-Net and Diffusion Transformer (DiT) decoder to capture a high level of detail and project each step’s noise. Furthermore, the model has an attention mechanism incorporated to capture global information and longer dependencies of the image input. This approach greatly boosts the quality of the image while keeping the textual information intact. To enhance the restoration further, the use of various losses enables the model to focus on the details of the image. In addition to conventional loss functions such as the L1 and L2, frequency loss and edge loss are also used. Together, these loss functions manage to keep the image’s texture, details, structures, and pixel-level accuracy. To assess its performance, the model is trained on the OLI2MSI dataset. Accordingly, the evaluation metrics demonstrate that the model can effectively reconstruct low-resolution images into high-resolution ones. The model achieved an FID (Fréchet Inception Distance) score of 11.47945, a PSNR (Peak Signal-to-Noise Ratio) score of 40.28188, an SSIM (Structural Similarity Index) score of 0.88802, and an LPIPS (Learned Perceptual Image Patch Similarity) score of 0.06426. The results confirm that the model outperforms the state of art techniques, in terms of numerical accuracy and perceptual quality, highlighting its usefulness for high-resolution remote sensing applications.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.