Artificial intelligence (AI) agents are rapidly evolving into autonomous systems capable of reasoning, acting, and continual evolving. Their increasing autonomy enables powerful real-world applications but also introduces security risks throughout the entire agent execution cycle. However, the security challenges arisi...
Dai-Zong Liu, Tian-Yao Luo, Shu-Wei Huang et al.· AI Plus· 0 citations
To achieve effective, stealthy, and persistent control, TrojanWorld combines Decision-Reflective Induction to steer trigger-conditioned imagination toward attacker-specified actions using decision feedback, Clean Behavior Anchoring to preserve trigger-free predictive and behavioral fidelity, and Causal Propagation to s...
Wen-Kai Huang, Si-Yuan Liang, Gaolei Li et al.· 0 citations
This work experimentally and systematically analyzes the differences between clean and jailbreak samples in the cross-attention feature space, revealing for the first time a cumulative separation effect and a progressively increasing trend of linear separability between the two during the diffusion process.
Si-Yuan Liang, Yupeng Qiu, Junfeng Fang et al.· 0 citations
This work develops a four-part, intent-oriented taxonomy that organizes multi-turn jailbreaks by adversarial intent structure and finds that effectiveness is driven by how deliberately intent is organized across turns rather than by context length or query count.
Siyuan Li, Aodu Wulianghai, Zehao Liu et al.· 0 citations
This work studies the channel delivering search and page observations is a fragile security boundary and introduces Authority-Chain Hijack (ACH), an expert-refined strategy that turns isolated search-result and page-content manipulations into a coherent evidence chain across seemingly corroborating sources.
Xuebin Li, Han-Qing Zhao, Si-Yuan Liang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.