This work develops a four-part, intent-oriented taxonomy that organizes multi-turn jailbreaks by adversarial intent structure and finds that effectiveness is driven by how deliberately intent is organized across turns rather than by context length or query count.
Siyuan Li, Aodu Wulianghai, Zehao Liu et al.· 0 citations
This paper presents a proactive defense framework for securing LLMs against evolving multi-turn adversarial attacks that combines disruption, misdirection, and adaptation across successive interaction turns and employs a cooperative multi-agent architecture.
Si-Yuan Li, Zehao Liu, Hao-Yu Li et al.· 0 citations
Safety Harness Evolution (SHE) is proposed, a framework that learns evolving safe boundaries from rollout trajectories and introduces an attribution-guided evolution loop that converts trajectory failures into structured diagnoses, learns artifact-specific boundary refinements, and selects evolved harnesses through safety-utility validation.
Wanying Qu, Qinghua Mao, Yu Li et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.