We show that aligned vision-language models also condition refusal on a property of a request's form: whether an image is attached, holding everything the request asks fixed. Attaching a blank canvas, an image that cannot be read, cannot relate to the request, and is byte-identical across every prompt in its condition,...
Haoyu Zhang, Yi Feng, Han-Wen Liu et al.· 0 citations
We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack success rate (ASR). The operative change is in the defense pipeline, not in the image. Across five...
Haoyu Zhang, Xiang-Chen Guan, Shi-Bo Zheng et al.· 0 citations
An empirical safety– utility ceiling for the non-iterative recovery-based defenses the authors evaluate, recurring across every guard and both target VLMs, is exposed.
Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain...
Haoyu Zhang, Zhuo-Xiang Wang, Shi-Bo Zheng et al.· 0 citations
A coding agent that installs packages and untangles version conflicts is implicitly reasoning about semantic-versioning constraints and dependency resolution. Whether current language models can actually do this has not been measured, and that is the gap we address. DepResolve-Bench is a programmatically generated benc...
Zhuo-Xi Wang, Haoyu Zhang, Jing-Wen Hou et al.· 2026 8th International Confe...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.