Skip to content
Open access

HARC-Net: Hierarchical Multiaxis Representation and Adaptive Residual Calibration for End-to-End SAR-to-Optical Image Translation

2026 · IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing · Vol 19, pp. 25871-25883 · 0 citations · 42 references

TL;DR

HARC-Net is presented, an end-to-end Transformer-based regression framework that combines hierarchical multiaxis representation learning with statistics-guided skip-feature calibration to improve robustness and reconstruction quality in remote sensing applications requiring geometrically consistent and noise-resilient optical reconstruction under adverse imaging conditions.

Abstract

Synthetic aperture radar (SAR) enables all-weather Earth observation; however, its inherent multiplicative speckle noise and geometry-dependent distortions pose significant challenges for SAR-to-optical image translation, often leading to structural deformation and degraded texture fidelity. To address these issues, this article presents Hierarchical MultiAxis Representation and Adaptive Residual Calibration Network (HARC-Net), an end-to-end Transformer-based regression framework that combines hierarchical multiaxis representation learning with statistics-guided skip-feature calibration to improve robustness and reconstruction quality. At the core of the proposed approach is a variable-axis sparse transformer (VASTormer) encoder, which integrates convolutional inductive bias with hierarchical multiaxis attention, including local block attention and sparse grid attention. This task-oriented encoder design enables efficient modeling of long-range dependencies while maintaining stable feature representations under speckle perturbations. To mitigate noise propagation in U-shaped architectures, we further introduce an adaptive dual attention and residual calibration (ADARC) module for skip connections. ADARC combines multistatistic spatial pooling (mean, max, min, and sum) with channelwise attention and learnable residual gating, effectively suppressing speckle-sensitive responses and improving semantic alignment between encoder and decoder features. Extensive experiments on two paired benchmarks, SEN1-2 and QXS-SAROPT, demonstrate that HARC-Net consistently achieves superior performance in both reconstruction quality and structural fidelity. The proposed method significantly reduces speckle-induced artifacts while preserving fine geometric details and linear structures. These results highlight the effectiveness of combining hierarchical local–global representation learning with statistics-guided feature calibration for robust cross-modal translation in remote sensing applications requiring geometrically consistent and noise-resilient optical reconstruction under adverse imaging conditions.

Read PDF

Similar papers

Open access Sep 2026

DSAN: Dual-Scale Aligned Network with Asymmetric Priors and Differentiable Soft-Edge Loss for SAR-to-Optical Image Translation

Synthetic aperture radar (SAR) provides all-weather imaging but faces challenges in visual interpretation due to low contrast and coherent speckle noise. To address contrast deficiency and edge blurring in SAR-to-optical translation, we propose a Dual-Scale Aligned Network (DSAN) built upon Pix2PixHD. First, an asymmetric dual-prior architecture is designed: the global generator ingests low-resolution SAR images enhanced by histogram equalization to capture macroscopic structures, while the local generator utilizes original high-resolution SAR images to preserve microscopic details, alleviating the trade-off between contrast and fine textures. Second, a Dual-Scale Fusion Module (DSFM) coupling Large Kernel Attention and Collaborative Attention breaks scale barriers, enabling bidirectional cross-scale alignment and deep fusion. Third, a continuous differentiable soft-edge loss is formulated using logarithmic dynamic range compression to prevent highlights from dominating gradients and enforce boundary consistency across urban areas, water bodies, and farmlands. Experiments on the Nanjing and public SEN1-2 datasets demonstrate that DSAN outperforms state-of-the-art models—including Pix2PixHD, CycleGAN, MSTMNet, and ICMA—in perceptual distribution realism (FID) with the sharpest geometric boundaries. Ablation studies confirm the effectiveness of the asymmetric dual-prior design, DSFM, and the refined edge loss.

Ying-Ying Kong, Dong-Ming Wang · 0 citations
#machine learning Preprint Aug 2026

FiLM-GPNet: Geometry-Aware Pseudo-Supervised Phase Restoration with Zero-Shot Generalization for Large Temporal InSAR Stacks

FiLM-GPNet is proposed, a geometry-conditioned network for wrapped-phase restoration that explicitly adapts to acquisition differences using Feature-wise Linear Modulation (FiLM) and a 7D per-pair geometry descriptor, supporting geometry-conditioned restoration as an effective alternative to fixed classical filtering across heterogeneous stacks.

Getnet Demil, Muhammad Farhan Humayun, Tomi Westerlund et al. · 0 citations
Preprint Sep 2026

Bridging Modalities and Tasks: A Unified Hierarchical ViT for SAR-to-Optical Translation and Semantic Segmentation

Synthetic Aperture Radar (SAR) images have all-weather, day-and-night observation capabilities. However, compared with optical images, their speckle noise and non-intuitive scattering mechanism limit the interpretability of the images. Generative models for SAR-to-optical (S2O) conversion can improve visual interpretability, but existing methods often ignore the constraints on semantic structure, which are necessary for downstream tasks, for the sake of visual effects. We propose a unified collaborative dual-task learning framework, termed BMT (Bridging Modalities and Tasks), that jointly optimizes S2O image translation and semantic segmentation through a shared hierarchical Vision Transformer. The framework integrates: (1) a LocalViTBlock that fuses global self-attention with spatial depthwise convolution through a learnable gating mechanism; (2) an enhanced output module combining multi-scale refinement processing, color correction and anti-aliasing, which calibrates channel-level color statistics through feature fusion; (3) a ControlNet-style conditional injection mechanism that encodes SAR wavelet features and segmentation labels into a multi-scale feature pyramid and injects them at each encoder layer through zero-initialized convolution; (4) a bounded Kendall uncertainty weighting scheme that prevents either task from dominating the shared representation. We evaluate the framework under both paired and unpaired translation settings, on the public WHU-OPT-SAR paired dataset and a self-constructed unpaired ship dataset built from HRSID and DIOR, respectively. The experimental results show that the proposed method achieves competitive S2O translation quality and semantic segmentation performance. The dataset and source code have been publicly released at https://github.com/Lewisyuaner/BMT-S2O-main.

Si-Yuan Liu, Xu-Ze Zhang, Yong-Shun Wang et al. · 0 citations
Open access Aug 2026

MTC-Net: Leveraging Multi-Temporal Consistency and Multi-View Synergistic Contrastive Learning for Remote Sensing Scene Classification

The remote sensing scene classification (RSSC) task plays a pivotal role in Earth observation missions, yet its progress remains constrained by the scarcity of high-quality labeled imagery. This article introduces a self-supervised learning (SSL) paradigm to address this challenge. First, for pseudo-label construction, a large set of long-interval satellite revisit imagery is collected and processed with pixel-level registration. The SIFT inliers retained during registration serve as saliency priors to guide asymmetric masking across views. This produces positive pairs that preserve global scene consistency while introducing controlled object-level ambiguities. Second, we propose a progressive layer-wise contrastive learning framework (MTC-Net) that couples the pseudo-label with the network’s representational hierarchy, forming a curriculum from local texture robustness to global semantic invariance. A dual-attention module with spatial–channel branches is further embedded to recalibrate intermediate features. The learning paradigm encourages the model to perform cross-view contextual reasoning rather than relying on pixel-wise correspondences. Experiments on three widely used datasets demonstrate that MTC-Net achieves competitive classification accuracy under limited-label settings, while ablation and visualization studies validate the effectiveness of establishing scene-level invariance through multi-temporal contrastive alignment.

Xiao Xiao, Han Zhang, Kenan Cheng et al. · 0 citations
Open access Jul 2026

ARGO-Net: An Adaptive Receptive-Field and Geometry-Oriented Network for Lightweight Ship Detection in Complex Maritime Scene

Deploying robust ship detectors in real-world maritime environments is severely bottlenecked by the dual challenges of strictly constrained computational resources and complex background interferences, such as dense berthing, wake patterns, and SAR speckle. To solve these problems, we propose ARGO-Net, a highly efficient architecture tailored for multi-modal maritime detection. At its core, ARGO-Net extracts physically meaningful and highly discriminative features through three targeted innovations. First, a Background Suppression and Reconstruction Module (BSRM) is developed to mitigate irregular coastal clutter and speckle in the frequency domain, reconstructing resilient spatial representations. Second, to capture the intrinsic morphological properties of ships, the High-Resolution Preserving Feature Network (HRPFN) employs geometry-oriented strip convolutions alongside an adaptive scale mechanism, effectively preserving the structural continuity of elongated hulls across extreme scale variations. Finally, a Semantic–Detail Alignment Fusion (SDAF) module is introduced to resolve cross-level spatial mismatches, ensuring that deep semantic context precisely informs low-level boundary localization. Extensive evaluations on the SeaShips and SSDD benchmarks highlight the exceptional efficiency–accuracy balance of ARGO-Net. With a marginal footprint of merely 2.3 M parameters and 6.9 G FLOPs, ARGO-Net achieves 97.9%/75.4% (mAP@50/mAP@50:95) on SeaShips and 99.5%/79.3% on SSDD. The proposed framework demonstrates that integrating background-aware feature reconstruction with geometry-driven fusion yields state-of-the-art localization precision without compromising lightweight deployability.

Jing Qu, Qiang Zhou, Bi-Meng Zhang et al. · 0 citations
Jul 2026

PriSAR: 3D Geometric-Prior-Guided Diffusion for Parameter-Controlled SAR Image Generation

The results support the conclusion that a lightweight 3D geometric prior improves viewpoint adherence for controllable SAR generation; it is intended as generation guidance rather than high-fidelity electromagnetic construction.

Fan Zhang, Xuanting Wu, Fei Ma et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.