A Review of Remote Sensing Image Translation Based on Generative Adversarial Networks
Abstract
Generative adversarial networks (GANs) have become an important approach for translating remote sensing images across sensing modalities, acquisition domains, and degraded observation conditions. This paper reviews 111 studies published between 2018 and 2026 and analyzes the field from three translation perspectives: intra-modal or domain translation, cross-modal translation, and image restoration or completion. The review also examines the use of translated imagery in change detection, object detection, semantic segmentation, and data augmentation. Existing methods are compared in terms of supervision strategy, data pairing, network architecture, feature representation, loss-function design, datasets, and evaluation practice. The literature shows a strong concentration on SAR-to-optical translation, accompanied by growing use of multiscale feature extraction, attention mechanisms, contrastive learning, cross-modal alignment, and pretrained representations. However, translation quality remains closely related to dataset characteristics, sensor configuration, preprocessing, and evaluation protocol. Pixel-level, structural, perceptual, distributional, sensor-related, and downstream metrics provide complementary evidence of translation performance. Major challenges include the non-unique nature of cross-modal mappings, preservation of geometric and sensor-related information, dependence on auxiliary observations in restoration tasks, and limited evidence for cross-domain generalization. Future research should develop task- and sensor-aware evaluation and strengthen cross-region and cross-sensor validation.