Artificial intelligence (AI) agents are rapidly evolving into autonomous systems capable of reasoning, acting, and continual evolving. Their increasing autonomy enables powerful real-world applications but also introduces security risks throughout the entire agent execution cycle. However, the security challenges arisi...
Dai-Zong Liu, Tian-Yao Luo, Shu-Wei Huang et al.· AI Plus· 0 citations
DARWIN, an evolutionary attack-defense framework that models jailbreaking as a continual process and updates guardrails through an attack-defense loop is proposed, an evolutionary attack-defense framework that models jailbreaking as a continual process and updates guardrails through an attack-defense loop.
DataShield is a data assessment framework that identifies risky fine-tuning samples and response segments through consensus subspace alignment over joint safety-critical semantic spaces derived from multiple safety-aligned LLMs, allowing both sample-level filtering and fine-grained segment-level masking.
Ze-Feng Wu, Weiwei Qi, Jielong Chen et al.· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.