Skip to content
Conference

FGSM Adversarial Example Generation Method Based on Target Region Constraint and Cross-Model Gradient Fusion

Jul 2026 · 2026 8th International Conference on Electronics and Communication, Network and Computer Technology (ECNCT) · pp. 287-294 · 0 citations · 13 references

Abstract

Deep learning-based object detection models are widely applied in remote sensing fields such as remote sensing image interpretation and maritime surveillance, yet their poor robustness against adversarial perturbations has become a critical security risk for the deployment of remote sensing vision systems. Existing FGSM-based adversarial example generation methods rely on gradients from a single model and apply perturbations across the entire image, suffering from insufficient cross-model attack transferability, low perturbation concealment and utilization efficiency, and thus failing to adapt to the characteristics of remote sensing images including large target scale variations and complex backgrounds. To address this, this paper proposes a target-region constrained cross-model FGSM adversarial example generation method for remote sensing images, which fuses the loss function gradients of YOLOv5 and RetinaNet to generate adversarial directions, and combines a region constraint mechanism based on detection box dilation to confine perturbations to targets and their surrounding areas. Experiments on the SSDD remote sensing ship dataset demonstrate that, compared with the traditional single-model FGSM, the proposed method significantly reduces the recall rate of remote sensing detection models while achieving better performance in the SSIM metric. It realizes an effective trade-off between attack intensity and visual concealment, verifying its feasibility and effectiveness in the remote sensing domain.

View source

Similar papers

2026

Dual-Domain Adversarial Purification for Robust Remote Sensing Scene Classification

Deep learning has boosted remote sensing (RS) scene classification, but adversarial examples can still cause high-confidence misclassification with imperceptible perturbations. Adversarial purification (AP) offers a practical test-time defense without retraining the classifier. However, most existing methods are confined to pixel-space restoration, which may leave residual adversarial effects that persist and amplify through feature extraction, ultimately biasing the prediction. To address these issues, a dual-domain AP (DDAP) framework is proposed to mitigate adversarial effects at both the pixel and feature levels in a unified pipeline. In the pixel domain, a pixel-domain frequency-aware diffusion purification (PFDP) module performs diffusion-based restoration through a frequency-aware dual-stream U-Net (FD-UNet). By integrating adaptive spectral filtering with multidomain consistency constraints, PFDP reduces adversarial-perturbation-dominated high-frequency responses while preserving structural details and semantic information in RS imagery. In the feature domain, an adversarial vulnerable channel dropout (AVCD) strategy models unshifted shallow-feature statistics with a Gaussian mixture model (GMM) and adaptively assigns channelwise dropout probabilities based on a samplewise shift score and channel vulnerability, thereby suppressing residual adversarial influence before downstream classification. Extensive experiments on UC Merced (UCM) and aerial image dataset (AID) across multiple backbones and attack types demonstrate that DDAP consistently improves robustness while maintaining a favorable clean–robust balance compared with representative baselines.

Yuru Su, Shaohui Mei, Mingyang Ma et al. · 0 citations
Preprint Aug 2026

ColorFD: A Finite-Difference Guided Black-Box Physical Adversarial Attack for Remote Sensing Object Detection

Although deep neural network-based remote sensing object detectors have achieved strong performance, they remain vulnerable to adversarial perturbations. Existing studies mainly focus on digital or white-box settings, whereas black-box physical attacks remain underexplored. These attacks are often constrained by limited physical feasibility and inefficient optimization in high-dimensional search spaces. To address these challenges, this paper proposes ColorFD, a black-box physical attack based on multiple pure-color patches. The patch positions and color parameters are jointly optimized using Differential Evolution (DE). A target-wise fitness and selection mechanism evaluates the attack state of each target and preserves target-specific improvements during evolution. Two guidance strategies further constrain the patch search space. Key-region localization identifies sensitive regions through finite-difference color probing. Common-feature extraction provides category-level spatial priors and avoids repeated localization. Although evaluated on aircraft, the formulation is not inherently restricted to this category. Experiments on YOLOv3u, YOLOv5u, and Faster R-CNN show that ColorFD outperforms the tested black-box patch method across all evaluated detectors and remains competitive with strong white-box baselines. Physical-world experiments further demonstrate that the optimized pure-color patches can be transferred from the digital domain to real imaging conditions.

Tian Guo, Guhang Qiu, Yuzhen Xie et al. · 0 citations
Jul 2026

GeoThreat: Transferable Targeted Adversarial Attacks on Large Vision-Language Models for Remote Sensing Image Interpretation

Adversarial attacks against large vision-language models (LVLMs) serve as an effective means of assessing their robustness in cross-modal semantic understanding. Existing studies mainly focus on corrupting visual inputs to induce predefined erroneous responses in general vision-language tasks, whereas corresponding investigations in remote sensing fields remain largely underexplored. Compared with natural image understanding, remote sensing image interpretation requires joint reasoning over local discriminative cues and global scene context. This poses additional challenges to achieving transferable semantic manipulation toward specified responses under black-box settings. To tackle these challenges, we propose GeoThreat, a transferable targeted adversarial attack method against LVLMs for remote sensing image interpretation. Specifically, GeoThreat modulates adversarial representations in accordance with the target content at both conceptual and perceptual levels. The class tokens from surrogate image encoders are employed as conceptual representations, while perceptual representations are distilled from patch tokens of the adversarial example through collaborative importance estimation. Beyond merely rolling out attention scores across layers, we incorporate adversarial-target similarity gradients to more faithfully characterize the relevance of local visual cues to the intended semantic manipulation. The perceptual representations are then dynamically aligned with target patch tokens in a cross-attentive manner, facilitating the adaptation of local cues toward designated semantic details. Finally, adversarial perturbations are iteratively updated via ensemble-based joint optimization of conceptual calibration and perceptual adaptation. Extensive experiments across diverse LVLMs demonstrate the superiority of GeoThreat in both transferability and controllability.

Yimin Fu, Yuefeng Bai, Baicheng Pan et al. · 0 citations
Review Aug 2026

A Comprehensive Review on Adversarial Attacks and Detection Techniques in Deep Learning Models for Image Analysis

The research methodology involved a systematic literature review using the Scopus database, adhering to Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines, and focusing on recent advancements in attack and defence techniques.

Reeti Jaswal, Vikas Khullar, Surya Narayan Panda · 0 citations
Sep 2026

Seeing Through Threats: (ADEx) Adversarial Detection through Explainability.

Deep Neural Networks (DNNs) remain vulnerable to adversarial perturbations, raising significant concerns in image processing applications, particularly in high-stakes domains such as medical imaging and security-critical systems. Most existing defense strategies are limited by domain specificity, architectural dependence, or the need for extensive retraining, making them impractical for real-world deployment. In this work, we propose ADEx, the first framework to integrate low-rank image approximation with explainability-driven analysis for the detection of adversarial samples. ADEx works by extracting a low-rank representation of the input image using Singular Value Thresholding (SVT), and identifying important image regions by computing class-specific gradient maps from the final layers of the classifier. These maps are then compared using Rank-Biased Overlap (RBO) to quantify the degree of attention drift induced by adversarial perturbations. ADEx is designed for adversarial detection in image classification systems, where class-specific gradient-based explanations are well defined. The framework operates without retraining or architectural modification and can be applied to a wide range of differentiable classifiers, provided gradient access is available for explanation generation. Extensive experiments across multiple datasets, architectures, and attack types demonstrate consistent performance, robustness to hyperparameter choices, and low sensitivity to calibration size. The method provides an interpretable and lightweight solution suitable for practical deployment.

Syamantak Sarkar, Nirmal Joseph, Sudhish N. George et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.