Supervised fine-tuning can substantially improve the downstream utility of large language models (LLMs) but may compromise their safety. Existing safety-preserving methods constrain downstream updates using safety-related parameters or subspaces, but mainly focus on safety preservation rather than joint safety and util...
Wei-Wei Qi, Chong-Yu Wang, Tian-Hang Zheng et al.· 0 citations
Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavioral safety and introduce ActBench, a self-evolving benchmark that evaluates such behavior risk from execution trajectories rather than final responses. Each case pairs a...
H. Yao, Yimin Liu, Meihui Chen et al.· 0 citations
DARWIN, an evolutionary attack-defense framework that models jailbreaking as a continual process and updates guardrails through an attack-defense loop is proposed, an evolutionary attack-defense framework that models jailbreaking as a continual process and updates guardrails through an attack-defense loop.
DataShield is a data assessment framework that identifies risky fine-tuning samples and response segments through consensus subspace alignment over joint safety-critical semantic spaces derived from multiple safety-aligned LLMs, allowing both sample-level filtering and fine-grained segment-level masking.
Ze-Feng Wu, Weiwei Qi, Jielong Chen et al.· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.