Skip to content
Preprint

CertVLA: Certified Defense against Physical Visual Attacks for Vision-Language-Action Models

Aug 2026 · 0 citations · 46 references
Computer Science

Abstract

Vision-Language-Action (VLA) policies are vulnerable to localized physical perturbations, yet existing certified patch defenses target discrete labels and cannot directly certify continuous, temporally correlated actions. We introduce CertVLA, a certified defense for closed-loop VLA control under bounded patch and texture attacks. CertVLA proposes a calibrated region of behaviorally consistent actions, while deterministic covering masks ensure that at least one checked prediction is attack-free. Specifically, CertVLA normalizes action disagreement by the benign variation of each mask pair and accepts a single-mask anchor only when it remains consistent under every second mask. It then calibrates the resulting max-min-max episode score to provide finite-sample clean coverage. Conjoining query-level decisions extends the action certificate to the complete closed-loop rollout. Furthermore, we prove that against any adaptive attacker satisfying the bounded-support threat model, every rollout certified by CertVLA executes only action chunks consistent with attack-erased clean predictions. Under dual-mask rollout correctness, this consistency certificate further guarantees task success. The certificate is independent of patch content, generation method, and physical transformation. Experiments in simulation and the real world demonstrate the empirical and certified effectiveness of CertVLA against patch attacks, with additional simulation validation on texture attacks.

View source

Similar papers

Open access Aug 2026

Sparse Adversarial Patch Attack and Robustness Evaluation Algorithm for Vision-Language Models

Visual language models (VLMs) have demonstrated outstanding performance in high-value domains such as autonomous driving, unmanned system navigation, and intelligent question-answering; however, the security of their cross-modal alignment mechanisms has not yet been fully verified. Existing visual adversarial patch att...

T.-Y. Chen, X.-Y. Hu, J.-F. Wang et al. · 0 citations
Preprint Sep 2026

Beyond Patch Removal: Persistent Adversarial Effects in Vision-Language-Action Policies

Adversarial patches to Vision-Language-Action (VLA) policies can cause both immediate action corruption and persistent state effects that remain after the patch is removed. Existing evaluations largely focus on continuous attacks and do not separate these two effects. We introduce a state-restoration protocol that remo...

Enjia Wu, Fu-Sen Guo, Yu-Xin Cao et al. · 0 citations
Preprint Sep 2026

Do Input-Level Defenses Transfer to Observation-Level Attacks on VideoLLMs?

Video Large Language Models (VideoLLMs) are increasingly deployed in safety-critical applications such as content moderation and video analytics. To process long videos efficiently, VideoLLMs rely on frame sampling, token compression, and modality fusion, which together form an observation pipeline that reduces the raw...

Bang-Shuo Zhu, Wei Song, Yu-Xin Cao et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models

Adversarial patches can disrupt Vision-Language-Action (VLA) models by manipulating visual observations, leading to failures in robot control. However, it remains poorly understood which internal mechanisms underlie these failures and how targeted interventions can mitigate them. In this work, we mechanistically analyz...

Yukiya Horiba, Koshiro Aoki, Shunsuke Yasuki et al. · 0 citations
Review Open access Sep 2026

Adversarial Robustness of Foundation Models for Intelligent Mechanical Systems: Threat Models, Benchmarks, and Defense Stacks

Foundation models increasingly operate across modalities (vision, language, audio, and vision–language) and are deployed in decision-critical pipelines with tool use and retrieval. This expands the adversarial surface: small perturbations to images or audio can flip predictions, carefully crafted text can induce unsafe...

Vishwanath · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.