Skip to content
Review Open access

Adversarial Robustness of Foundation Models for Intelligent Mechanical Systems: Threat Models, Benchmarks, and Defense Stacks

Sep 2026 · International Journal For Multidisciplinary Research · 0 citations · 24 references

Abstract

Foundation models increasingly operate across modalities (vision, language, audio, and vision–language) and are deployed in decision-critical pipelines with tool use and retrieval. This expands the adversarial surface: small perturbations to images or audio can flip predictions, carefully crafted text can induce unsafe actions, and cross-modal attacks can exploit representation alignment to produce consistent but wrong outputs. This paper reviews adversarial robustness of foundation models across modalities and proposes a unified benchmark-and-defense stack. We first formalize multimodal threat models (white-box/black-box, digital/physical, prompt-level/system-level) and show how attack objectives differ across classification, retrieval, captioning, and agentic planning. We then summarize benchmark families for robustness: standardized perturbation budgets in vision, imperceptible audio attacks, instruction-following adversarial prompts in language, and cross-modal attacks on vision–language alignment and retrieval. Finally, we present a practical defense stack combining (i) robust training and regularization, (ii) multimodal input sanitization and consistency checks, (iii) retrieval/verification and ensemble critics, and (iv) runtime guardrails for tool execution. We recommend reporting both utility and security metrics: clean accuracy, robust accuracy, attack success rate, confidence calibration, and worst-case safety violations under adaptive adversaries. The goal is to provide a deployment-oriented roadmap for measuring and improving robustness of multimodal foundation models.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.