Jul 2026· International Conference on Machine Vision and Applications· Vol 14270, pp. 1427003 - 1427003-13· 0 citations· 45 references
Engineering
TL;DR
It is established that architectural choice critically impacts adversarial robustness and that diffusion-based purification provides practical test-time defense for deployed autonomous driving systems.
Abstract
Traffic sign recognition systems are critical for autonomous vehicle safety yet remain vulnerable to adversarial perturbations that exploit environmental conditions. Existing naturalistic attacks employ simplified approximations without modeling underlying physics, while defenses lack comprehensive evaluation across diverse architectures and attack types. This work formulates six physically-grounded adversarial attacks spanning naturalistic perturbations (shadow, light patch, obstruction) and weather conditions (fog, snow, frost), optimized via Prior-guided Bayesian Optimization for black-box scenarios. We evaluate nine architectures across three geographically diverse datasets, revealing that transformer-based models exhibit 20.0 percentage points (pp) lower average attack success rate than CNNs (41.5% vs 61.5%) with comparable clean accuracy across datasets. We introduce Robustness Score (RS) to quantify resilience across all attacks, with baseline ConvNeXt-Tiny achieving 60.9% RS compared to best CNN at 50.7%. DiffPure diffusion-based purification substantially improves robustness through test-time noise injection and reverse denoising, increasing RS by 29.8-37.3 pp. Post-defense, transformer architectures achieve 6-16% residual attack success rate compared to 15-39% for CNNs, with consistent effectiveness across datasets (variance ⪅ 2.1 pp). These findings establish that architectural choice critically impacts adversarial robustness and that diffusion-based purification provides practical test-time defense for deployed autonomous driving systems.
Evaluated on GTSRB and LISA across four backbones and three physical attack types, LAMDA is the only method among ten evaluated that consistently improves robustness across all attack-backbone-dataset combinations, while preserving or improving clean accuracy in nearly all cases.
Pedram MohajerAnsari, Amir Salarpour, Mert D. Pesé· 0 citations
TRADES-JR, a TRADES loss function guided by Jacobian regularization, is proposed, which enables LNNs to maintain robust and high-accuracy traffic sign recognition even in adversarial environments, thereby enhancing the reliability of the autonomous driving system.
YunKai Zhao, Shang Gao, Jieliang Zhao· Proceedings of the Instituti...· 0 citations
Occupancy detection is fundamental to the operational intelligence of smart buildings, driving critical functions in energy management, HVAC automation, and physical security. While modern Deep Learning (DL) models have achieved high accuracy in parsing complex environmental sensor data, they remain highly vulnerable to adversarial examples, imperceptibly perturbed inputs designed to deceive neural networks. These vulnerabilities pose severe real-world risks, ranging from energy sabotage, in which systems heat empty rooms, to critical security breaches in which intruders go undetected. To address this security gap, we propose ADS-Guard, a novel Adversarial Detection and Sanitization (ADS) framework rooted in sequence-to-sequence autoencoder purification. Unlike standard denoising techniques, ADS-Guard incorporates a latent consistency regularization mechanism that encourages alignment between clean and adversarial representations in the latent feature space. We evaluated ADS-Guard using a comprehensive experimental pipeline comprising five distinct DL architectures (LSTM, GRU, 1D-CNN, MLP, and Transformer) across three diverse datasets: (1) The UCI Occupancy dataset (20,699 samples) for standard binary detection; (2) Building59 dataset (7200 samples) for three-class occupancy-level classification (Low, Medium, High); and (3) Room Occupancy dataset (10,129 samples), representing a highly imbalanced binary occupancy-detection task. We evaluate ADS-Guard against both Fast Gradient Sign Method (FGSM) and Projected Gradient Descent (PGD) attacks across diverse occupancy datasets and model architectures. We further assess the framework under adaptive white-box attacks and compare its performance with FGSM-based and PGD-based adversarial training baselines. Our results demonstrate that adversarial attacks can substantially degrade occupancy-detection performance across datasets and model architectures. ADS-Guard consistently improves robustness relative to undefended models against both FGSM and PGD attacks, recovering a substantial portion of the lost performance in binary occupancy tasks and providing meaningful gains in the more challenging multi-class setting. Furthermore, ADS-Guard remains effective under stronger adaptive threat models while providing a practical retraining-free defense that can be integrated with existing occupancy-detection systems without modifying downstream classifiers.
Pratiksha Chaudhari, Yang Xiao, Wei Sun· Italian National Conference...· 0 citations
This work reframe robustness as a detection problem, introducing a learned physics-informed detector whose output is fed to a hardened forecaster as an input feature and trained against adaptive attacks with the forecaster fixed and improves even on adversarial training hardened against the physics-aware attack itself.
Visual language models (VLMs) have demonstrated outstanding performance in high-value domains such as autonomous driving, unmanned system navigation, and intelligent question-answering; however, the security of their cross-modal alignment mechanisms has not yet been fully verified. Existing visual adversarial patch attacks typically rely on continuous, dense pixel perturbations, which are easily detected and blocked by anomaly detection systems in practical engineering applications. This paper proposes a novel sparse adversarial patch attack algorithm (Sparse Patch Attack, SPA), which successfully misleads the text generation results of VLMs by generating highly dispersed discrete pixel perturbations in non-salient regions of the image. To achieve this, we introduce a differentiable L0-norm approximation and a cross-attention masking mechanism to minimize the number of modified pixels. Furthermore, addressing the characteristics of large-scale model open-ended text generation, we construct a multi-dimensional robustness evaluation framework covering semantic deviation, target achievement rate, and visual concealment. Preliminary experiments on mainstream visual-language models (such as LLaVA and BLIP-2) demonstrate that the SPA algorithm can achieve high success rates in targeted cross-modal attacks with an extremely low pixel modification rate (<1%). This study reveals a novel security vulnerability in visual-language models within complex real-world environments and provides a quantitative evaluation benchmark for future defense mechanisms in multimodal models.
T.-Y. Chen, X.-Y. Hu, J.-F. Wang et al.· Advanced Electromagnetics· 0 citations
This research investigates the adversarial robustness of lane detection for Autonomous Vehicles (AVs) under challenging driving conditions using Generative Adversarial Networks (GANs). In this work, the term adversarial refers to the adversarial training mechanism of GANs and to robustness under naturally adverse driving conditions, particularly illumination variation, rather than to defence against deliberate pixel-level perturbation attacks such as FGSM or PGD. Lane detection is a crucial component for safe navigation, but it often fails under poor lighting or adverse weather. To solve this, a U-Net model is trained on the Berkeley DeepDrive (BDD100K) dataset as a baseline. Then, Conditional GAN (CGAN) is used with the Cityscapes dataset to learn the mapping between RGB images and lane masks, which improves structural consistency. To handle illumination changes, CycleGAN is used to simulate Day-to-Night and Night-to-Day translations using BDD100K datasets, creating a more diverse training set. Preprocessing involves resizing images to 512×512 to ensure training efficiency on limited GPU hardware. The experiments are conducted using TensorFlow in a GPU-accelerated environment. Results show that the U-Net + CycleGAN model achieves a Precision of 65.41% and an F1-Score of 63.79%, which outperforms previous studies. The CGAN model also shows high performance with 92.55% F1-Score. This research proves that using GANs for data augmentation and domain translation can enhance the adversarial robustness and reliability of lane detection systems in real-world scenarios.
Brian Lee Chong Ming, Thinesh Ganesan· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.