Haoyu Zhang, Yi Feng, Shi-Bo Zheng et al.
· 0 citations
Save
{ copied = true; setTimeout(() => copied = false, 1500) })"
class="icon-btn" aria-label="Copy link">
{ copied = 'apa'; setTimeout(() => { copied = null; open = false }, 1000) })"
class="flex w-full items-center justify-between rounded-lg px-3 py-2 text-left text-sm hover:bg-gray-100 dark:hover:bg-ink-800">
Copy APA
Copied ✓
{ copied = 'mla'; setTimeout(() => { copied = null; open = false }, 1000) })"
class="flex w-full items-center justify-between rounded-lg px-3 py-2 text-left text-sm hover:bg-gray-100 dark:hover:bg-ink-800">
Copy MLA
Copied ✓
{ copied = 'bibtex'; setTimeout(() => { copied = null; open = false }, 1000) })"
class="flex w-full items-center justify-between rounded-lg px-3 py-2 text-left text-sm hover:bg-gray-100 dark:hover:bg-ink-800">
Copy BibTeX
Copied ✓
We show that aligned vision-language models also condition refusal on a property of a request's form: whether an image is attached, holding everything the request asks fixed. Attaching a blank canvas, an image that cannot be read, cannot relate to the request, and is byte-identical across every prompt in its condition,...
Haoyu Zhang, Yi Feng, Han-Wen Liu et al.
· 0 citations
Save
{ copied = true; setTimeout(() => copied = false, 1500) })"
class="icon-btn" aria-label="Copy link">
{ copied = 'apa'; setTimeout(() => { copied = null; open = false }, 1000) })"
class="flex w-full items-center justify-between rounded-lg px-3 py-2 text-left text-sm hover:bg-gray-100 dark:hover:bg-ink-800">
Copy APA
Copied ✓
{ copied = 'mla'; setTimeout(() => { copied = null; open = false }, 1000) })"
class="flex w-full items-center justify-between rounded-lg px-3 py-2 text-left text-sm hover:bg-gray-100 dark:hover:bg-ink-800">
Copy MLA
Copied ✓
{ copied = 'bibtex'; setTimeout(() => { copied = null; open = false }, 1000) })"
class="flex w-full items-center justify-between rounded-lg px-3 py-2 text-left text-sm hover:bg-gray-100 dark:hover:bg-ink-800">
Copy BibTeX
Copied ✓
2026
An empirical safety– utility ceiling for the non-iterative recovery-based defenses the authors evaluate, recurring across every guard and both target VLMs, is exposed.
Haoyu Zhang, Zhuo-Xiang Wang, Shi-Bo Zheng et al.
· arXiv.org · 0 citations
Save
{ copied = true; setTimeout(() => copied = false, 1500) })"
class="icon-btn" aria-label="Copy link">
{ copied = 'apa'; setTimeout(() => { copied = null; open = false }, 1000) })"
class="flex w-full items-center justify-between rounded-lg px-3 py-2 text-left text-sm hover:bg-gray-100 dark:hover:bg-ink-800">
Copy APA
Copied ✓
{ copied = 'mla'; setTimeout(() => { copied = null; open = false }, 1000) })"
class="flex w-full items-center justify-between rounded-lg px-3 py-2 text-left text-sm hover:bg-gray-100 dark:hover:bg-ink-800">
Copy MLA
Copied ✓
{ copied = 'bibtex'; setTimeout(() => { copied = null; open = false }, 1000) })"
class="flex w-full items-center justify-between rounded-lg px-3 py-2 text-left text-sm hover:bg-gray-100 dark:hover:bg-ink-800">
Copy BibTeX
Copied ✓
Preprint
Jul 2026
Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain...
Haoyu Zhang, Zhuo-Xiang Wang, Shi-Bo Zheng et al.
· 0 citations
Save
{ copied = true; setTimeout(() => copied = false, 1500) })"
class="icon-btn" aria-label="Copy link">
{ copied = 'apa'; setTimeout(() => { copied = null; open = false }, 1000) })"
class="flex w-full items-center justify-between rounded-lg px-3 py-2 text-left text-sm hover:bg-gray-100 dark:hover:bg-ink-800">
Copy APA
Copied ✓
{ copied = 'mla'; setTimeout(() => { copied = null; open = false }, 1000) })"
class="flex w-full items-center justify-between rounded-lg px-3 py-2 text-left text-sm hover:bg-gray-100 dark:hover:bg-ink-800">
Copy MLA
Copied ✓
{ copied = 'bibtex'; setTimeout(() => { copied = null; open = false }, 1000) })"
class="flex w-full items-center justify-between rounded-lg px-3 py-2 text-left text-sm hover:bg-gray-100 dark:hover:bg-ink-800">
Copy BibTeX
Copied ✓