Large language models (LLMs) remain highly vulnerable to jailbreak attacks that induce harmful behaviors and circumvent safety alignment. To defend against such attacks, adversarial training paradigms have been proposed to first simulate failure modes and then train the model to correct them, yielding promising improve...
Luo-Yu Chen, Wei-Qi Wang, Chen-Han Zhang et al.· 0 citations
Large language models (LLMs) remain vulnerable to jailbreak attacks that conceal harmful intent within complex adversarial prompts. Existing defenses primarily rely on input perturbation or harmful-output suppression, but they rarely model where malicious intent resides, resulting in brittle protection and excessive ov...
Luo-Yu Chen, Wei-Qi Wang, Chen-Han Zhang et al.· 0 citations
CUNO is proposed, a curriculum-based graph unlearning framework that removes the forget set progressively, ordering samples by their estimated unlearning difficulty across multiple stages, and employs a distribution-level negative preference optimization objective at each curriculum stage that steers the model away fro...
Chenhan Zhang, Ali Braytee, M. Bandara et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.