Skip to content

Similar papers

#artificial intelligence Preprint Aug 2026

The Uncontrolled Variable: Vision-Language Refusal Is Conditioned on the Image-Attachment Interface, and Not Robust to Irrelevant Image Properties

We show that aligned vision-language models also condition refusal on a property of a request's form: whether an image is attached, holding everything the request asks fixed. Attaching a blank canvas, an image that cannot be read, cannot relate to the request, and is byte-identical across every prompt in its condition,...

Haoyu Zhang, Yi Feng, Han-Wen Liu et al. · 0 citations
Preprint Aug 2026

TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint

When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internally distinguish when abstention is required, but fail to express it anyway. We introduce TRAPSBench, a procedurally generated video benchmark of 1,404 matched physics pairs in which a...

Fnu Pramono, J. Cai, Sourabh Kulkarni · 3 citations · ⚡1
#artificial intelligence Preprint Sep 2026

Knowing When Not to Answer: Abstention and Refusal Reasoning in Vision--Language Models

Many medical conditions require diagnosis through detailed, multi-context clinical assessment rather than from visual appearance alone. Despite this, vision-language models (VLMs) are increasingly queried to interpret images in ways that touch on medical or diagnostic judgments, raising safety concerns when such infere...

Karan Dua, Amit Agarwal, Hitesh Laxmichand Patel et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection Guardrails

Verdict-only evaluation does not reveal whether a vision-language model (VLM) used the visual evidence that should support its decision. We study this problem in web-agent guardrails, where a VLM judges whether on-screen text conflicts with a user instruction. We introduce Mind2Web-Injection, a benchmark of 9,954 instr...

Suyoung Lee, Myungsub Choi · 0 citations
#artificial intelligence Preprint Sep 2026

Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models

Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding compliance with requests that are incorrect, unsafe, infeasible, or unanswerable. However, existing benchmarks predominantly evaluate non-compliance at the level of the query as a whole, assuming that each request...

Minji Kim, Jihyoung Jang, Hyounghun Kim · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.