Supervised fine-tuning can substantially improve the downstream utility of large language models (LLMs) but may compromise their safety. Existing safety-preserving methods constrain downstream updates using safety-related parameters or subspaces, but mainly focus on safety preservation rather than joint safety and util...
Wei-Wei Qi, Chong-Yu Wang, Tian-Hang Zheng et al.· 0 citations
DARWIN, an evolutionary attack-defense framework that models jailbreaking as a continual process and updates guardrails through an attack-defense loop is proposed, an evolutionary attack-defense framework that models jailbreaking as a continual process and updates guardrails through an attack-defense loop.
This work introduces \textsc{JailbreakSkill}, a skill-centric framework for scaling automated red-teaming through reusable and continuously evolving attack capabilities, which packages existing attack strategies into modular, agent-ready skills that can be directly reused and adaptively selected across tasks and target...
Xiaoyu Wen, Jiajia Li, Zhida He et al.· 2 citations
This work proposes the first completely probabilistic architecture-agnostic guardrail to leverage the LLM early output distributional signals for estimating and calibrating the safety probability, thereby enabling early stopping of unsafe ongoing outputs.
Xinzhe Huang, Biwu Yao, Kedong Xiu et al.· 0 citations
A Unidirectional Safety Gate (USG) is proposed, instantiated as a Null Space Cubic Layer together with an Inverse Adapter inserted after the final Transformer layer, suggesting that release-time representation-space blocking can raise the cost of malicious downstream adaptation without requiring downstream cooperation.
Yuxuan Huang, Xingyu Zeng, Tian-Hang Zheng et al.· 0 citations
This work introduces AgentSnare, a trajectory-adaptive deception system that dynamically unfolds a decoy environment to continually steer the penetration agent away from the real target.
Ruoyu Wang, Heng Zhao, Renjie Wu et al.· 2 citations
DataShield is a data assessment framework that identifies risky fine-tuning samples and response segments through consensus subspace alignment over joint safety-critical semantic spaces derived from multiple safety-aligned LLMs, allowing both sample-level filtering and fine-grained segment-level masking.
Ze-Feng Wu, Weiwei Qi, Jielong Chen et al.· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.