Back to feed

Unmasking Toxicity and Vulnerabilities in Large Vision-Language Models

Jun 2026 · ACM Transactions on Intelligent Systems and Technology · 0 citations · 136 references

TL;DR

This study systematically examines the vulnerabilities of open and closed-weight LVLMs, including LLaVA, InstructBLIP, Fuyu, Qwen, DeepSeek, Gemini, GPT, and Grok, using adversarial prompting strategies informed by social theories to simulate real-world social manipulation tactics.

Abstract

The rapid advancement of Large Vision-Language Models (LVLMs) has enhanced their capabilities from content creation to productivity enhancement. Despite their innovative potential, LVLMs exhibit vulnerabilities, especially in generating potentially toxic or unsafe responses. Malicious actors can exploit these vulnerabilities to propagate toxic content using strategically crafted prompts without fine-tuning or compute-intensive procedures. Despite ongoing red-teaming efforts to identify and mitigate these risks, the exploration of LVLM vulnerabilities remains nascent and yet to be fully addressed in a systematic approach. This study systematically examines the vulnerabilities of open and closed-weight LVLMs, including LLaVA, InstructBLIP, Fuyu, Qwen, DeepSeek, Gemini, GPT, and Grok, using adversarial prompting strategies informed by social theories to simulate real-world social manipulation tactics. Our findings show that (i) toxicity and insult are the most prevalent behaviors, with mean toxicity scores 19.32% and 12.36%, respectively; (ii) Gemini-2.0-Flash, LLaVA-v1.6-Vicuna-13B, and Grok-2-Vision-1212 are the most vulnerable models. Their toxic response rates reach 46.93%, 23.81%, and 17.98%, respectively, while insult response rates reach 47.94%, 14.62%, 12.27%, respectively; (iii) prompting strategies incorporating dark humor and multimodal toxic prompt completion significantly elevate these vulnerabilities. Despite extensive safety alignment efforts, models still generate content with varying degrees of toxicity when prompted with adversarial inputs, highlighting the urgent need for enhanced safety mechanisms and robust guardrails in LVLM development.

View source

Similar papers

Review Open access 2026

Survey on Adversarial Prompt Generation and Robustness Analysis in Large Language Models

This survey provides a comprehensive analysis of adversarial prompting strategies, ranging from input manipulation techniques to semantic and structural distortions, and explores defense strategies across preprocessing, model-level, postprocessing, and hybrid strategies, highlighting recent advances and their limitations.

A. Nasution, Ahmet Emre Ergün, Aytu˘g Onan et al. · 0 citations
Preprint Aug 2026

Generating Attacks for LLMs with GFlowNets

This study proposes an automated, human-independent, and adaptive approach leveraging GFlowNets to identify LLM vulnerabilities by utilizing one large language model to test another, and introduces a model capable of generating attack inputs in the Turkish language.

Berkay Ozcam, Irem Onen, M. Amasyalı et al. · 0 citations

Understanding and Exploiting Phase Sensitivity for Attacking Large Vision–Language Models

This paper proposes a novel LVLM attack method, called BadPhase with further backdoor designs, to implant adversarial phase as triggers into any image inputs via data poisoning so as to control the LVLMs’ predictions and finds that LVLMs are sensitive to the phase-aware image structure.

Daizong Liu, Junhao Dong, Xiang Fang et al. · 0 citations
Conference Jul 2026

Generating Attacks for LLM with GFlowNets

The rapid advancement of Large Language Models (LLMs) has facilitated their ubiquitous integration into various domains, leading to widespread adoption. However, this escalating trend has introduced significant security vulnerabilities, necessitating the identification and mitigation of flaws arising from malicious exploitation. Red teaming assessments, conducted to evaluate model robustness through diverse adversarial inputs, are essential for exposing security risks and implementing countermeasures. Currently, red teaming is performed either manually by experts or automatically using predefined attack datasets. Nevertheless, manual testing remains time-consuming, while existing automated methods suffer from limited creativity due to their inherent dependency on fixed datasets. In this study, we propose an automated, human-independent, and adaptive approach leveraging GFlowNets to identify LLM vulnerabilities by utilizing one large language model to test another. Within this framework, an attacker model is trained against a specified victim model to perform automated red teaming and provide a quantitative robustness score. This research aims to generate more effective adversarial attacks in English compared to existing benchmarks and, as a novel contribution to the literature, introduces a model capable of generating attack inputs in the Turkish language.

Berkay Özçam, İrem Önen, E. I. Tatli et al. · 0 citations
2025

Towards Building Model/Prompt-Transferable Attackers against Large Vision-Language Models

A new perspective of information theory is introduced to investigate LVLMs’ transferable characteristics by exploring the relative dependence between outputs of the LVLM model and input adversarial samples and formulate the complicated calculation of information gain as an estimation problem and incorporate such informative constraints into the adversarial learning process.

Xiaowen Cai, Daizong Liu, Xiaoye Qu et al. · 7 citations