This paper introduces the concept of instruction-dense visual jailbreaks, in which image-generation models produce detailed, readable, and actionable harmful instructions within images, and proposes TYPO, a black-box framework that exploits this safety gap by automatically generating adversarial TYPOgraphy prompts.
Meng Xie, Li Zeng, Hang Zhang et al.· arXiv.org· 0 citations
CloakDiff is proposed, the first framework for reversible, high fidelity privacy protection against text-based query attacks in VLMs and EDM Heuristic Sampling, a principled diffusion schedule for adversarial guidance.
Qinghua Lu, Ziqi Zhou, Yufei Song et al.· 0 citations
The proposed LLM agent early-stopping cascade outperforms the best single-gate baseline in every model-environment pair, saving 1.5-8.8 times more compute at a 90% recall target.
Kai Ruan, Zihe Huang, Ziqi Zhou et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.