Prompt injection is a leading security risk for LLMs and LLM-based applications such as agents. State-of-the-art red-teaming methods for prompt injection leverage reinforcement learning (RL) to train an attacker LLM to generate effective injected prompts. However, when targeting frontier LLMs such as GPT-6-Luna, a majo...
Chen-Long Yin, Xiao-Long Jin, Wei Zou et al.· 0 citations
On-policy distillation (OPD) trains a student to match a teacher's next-token distributions on student-generated trajectories. However, privileged information supplied to the teacher for OPD training may contain sensitive data. Whether the student leaks private information about the records supplied to the teacher duri...
Zhexi Lu, Mingzhi Zhu, Stacy Patterson et al.· 0 citations
AnchorMixGAN is a generative semi-supervised framework that addresses distributed denial-of-service (DDoS) attacks in cloud-integrated IoT networks through anchor-aligned target construction, and analyzes how reference-classifier error, view construction, and sharpening affect the target.
Jin Yang, Xu-Feng Liu, Yong-Cai Hu et al.· 0 citations
We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifier evaluates 't Hooft anomaly matching, superpotential R-charge consistency, central-charge matching, and a bounded chiral-ring proxy. A claim that passes receives a consistency cert...
Xingyang Yu· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This paper proposes TRACE, which simulates harmful SFT trajectories to produce progressively corrupted model states, and optimizes a safety patch simultaneously across these states, and improves the safety rate by up to 77 percentage points over the best baselines, while maintaining comparable utility to the undefended...
AutoDojo is introduced, a generative benchmark built on AgentDojo and AgentDyn that generates IPI adaptively for a given agent and defense, and it is demonstrated that standard static benchmarks often significantly overestimate defense efficacy.
Xinhang Ma, Taoran Li, Chao-Wei Xiao et al.· 5 citations
To study this phenomenon, SocioHack is introduced, a sandbox of 72 societal environments, and it is found that within these environments, reward hacking naturally emerges and leads to regulatory loophole discovery.
Wei Liu, Xinyi Mou, Hanqi Yan et al.· arXiv.org· 3 citations
MemPoison is proposed, a novel memory poisoning attack that bypasses selective memory mechanisms in LLM agents, where an attacker can inject triggerable backdoors into the agent's long-term memory through dialogue interactions, thereby misleading its subsequent responses.
Hongtao Wang, Sean Yang, Yu Chen et al.· 5 citations
The Infinite Impostor is introduced, an attack model in which an autonomous agent interposes itself between two parties who already trust each other, hijacking an existing relationship rather than building a new one from scratch, and the governance tensions that arise when platforms become the regulatory substrate of d...
Osama Zafar, Alexander Nemecek, Erman Ayday· arXiv.org· 0 citations
AI agents increasingly act autonomously in the world, yet harmful behavior cannot be reliably traced to the account that deployed the agent. This creates an accountability gap across both benign and malicious settings: misconfigured or hijacked agents may cause unintended harm, while malicious operators may deploy agen...
Ruben Chocron, Doron Jonathan Ben Chayim, Eyal Lenga et al.· 0 citations
It is shown that even zero-bit watermarking supports internal attribution under per-entity multi-key deployments without explicitly encoding identity, and external identification in selected text and image configurations is demonstrated.
Toluwani Aremu, Nils Lukas, Jie Zhang· arXiv.org· 1 citation