2026· Journal of emerging investigators· 0 citations
TL;DR
This study set up a baseline image classifier for images of digits and attacked the images using the fast gradient sign method, and hypothesized that introducing adversarial training for this classifier would significantly improve downstream classification accuracy in all three algorithmic defense settings.
Abstract
An adversarial attack is a modification to the pixels of an image for the purpose of making a machine learning system misclassify the image. The foremost defense against adversarial attacks is adversarial training: a process in which the machine learning system trains on the already attacked images. But, this is not the only kind of defense. There are also algorithmic defense methods, which work to modify the learning process to be resilient to adversarial attacks without involving attacked examples. In this study, we considered three algorithmic defense settings: no algorithmic defense, defensive distillation, and gradient masking. Then we evaluated the role of adversarial training as part of defending machine learning models from adversarial attacks. Specifically, we set up a baseline image classifier for images of digits (MNIST dataset) and attacked the images using the fast gradient sign method. We hypothesized that introducing adversarial training for this classifier would significantly improve downstream classification accuracy in all three algorithmic defense settings. We found that, for all algorithmic defense settings applied to neural networks with between one and six convolutional layers, adding adversarial training consistently resulted in a statistically significant increase in accuracy. While these findings are limited by the specific data, parameters, and algorithms explored, our results suggest that implementing adversarial training within all lines of defense against adversarial attacks would be beneficial. We believe that this insight increases awareness of cybersecurity threats such as adversarial attacks and lowers the barrier of entry to protect against them.
The research methodology involved a systematic literature review using the Scopus database, adhering to Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines, and focusing on recent advancements in attack and defence techniques.
This paper introduces DefendMal, a novel framework that synergistically combines Denoise Autoencoder with Sequence Squeezing, a Context-aware Adversarial Generator (CAG-AdvGAN), Projected Gradient Descent (PGD) adversarial training, and a Positive–Negative Detector with Variational Autoencoder (PNDetector-VAE) to enhance robustness against evolving adversarial threats.
Dennis Benedict Crasta, Vikash Kumar· Journal of Computer Virology...· 0 citations
This study evaluated the adversarial robustness of machine learning-based fraud detection systems by comparing classifier vulnerability profiles and assessing adversarial training as a mitigation strategy. Using the IEEE-CIS Fraud Detection dataset, comprising 590,540 transactions with a fraud incidence of 3.5%, four classifiers—logistic regression, random forest, gradient boosting, and a feed-forward neural network—were trained under identical preprocessing and class-weighting conditions and then subjected to Fast Gradient Sign Method and Projected Gradient Descent attacks at a perturbation budget of 0.02. Adversarial examples were constructed directly using closed-form and backpropagated gradients for the differentiable classifiers and using a logistic regression surrogate for the non-differentiable ensembles, before adversarial training was applied as a post-attack mitigation stage. Logistic regression proved the most adversarially vulnerable architecture, sustaining a 31.12-percentage-point recall loss under Projected Gradient Descent, while adversarial training subsequently restored its recall from 0.39 to 0.999 at an accuracy cost of 0.10 percentage points. Random forest and gradient boosting were not degraded by the surrogate-based attack, indicating that comparative robustness claims for tree-based ensembles require attack methods suited to their non-differentiable structure rather than transfer-based evaluation alone. Within the scope of this single-dataset evaluation, the findings support the adoption of adversarial training for gradient-based fraud detection models and suggest that robustness claims should be accompanied by disclosure of the attack methodology used to establish them.
Ololade Zainab Adesokan, Abiola Omolola Bamsa, O. Obioha-Val et al.· Journal of Engineering Resea...· 0 citations
This paper advocates for a forward-thinking approach that balances technical sophistication with human-centric principles, ensuring that adversarial deep learning evolves into a discipline not just of technical defense, but also of trust, transparency, and accountability.
Maisam Abbas, Ran-Zan Wang· IEEE Open Journal of the Com...· 0 citations
The study systematically compares two distinct adversarial training strategies: ‘pre-train’, where adversarial examples are generated beforehand, and ‘in-train’, where perturbations are introduced dynamically during the training process, to understand the advantages and limitations of each approach in enhancing model robustness.
José María Jorquera Valero, Ibon Bengoechea Cazorla, Manuel Gil Pérez· IEEE Access· 0 citations
A resilient hybrid defense mechanism aimed to mitigate the impact of two potent adversarial attacks: Fast Gradient Sign Method (FGSM) and Carlini&Wagner (C&W) attack is developed.
Khushnaseeb Roshan· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.