Skip to content

Category

cybersecurity

1,065 papers

#machine learning Preprint Sep 2026

Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning

Prompt injection is a leading security risk for LLMs and LLM-based applications such as agents. State-of-the-art red-teaming methods for prompt injection leverage reinforcement learning (RL) to train an attacker LLM to generate effective injected prompts. However, when targeting frontier LLMs such as GPT-6-Luna, a majo...

Chen-Long Yin, Xiao-Long Jin, Wei Zou et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Leaky Students: Membership Inference against On-Policy Distillation

On-policy distillation (OPD) trains a student to match a teacher's next-token distributions on student-generated trajectories. However, privileged information supplied to the teacher for OPD training may contain sensitive data. Whether the student leaks private information about the records supplied to the teacher duri...

Zhexi Lu, Mingzhi Zhu, Stacy Patterson et al. · 0 citations
#machine learning Preprint Sep 2026

AnchorMixGAN: Anchor-Aligned Generative Semi-Supervision for DDoS Detection in Cloud-Integrated IoT Networks

AnchorMixGAN is a generative semi-supervised framework that addresses distributed denial-of-service (DDoS) attacks in cloud-integrated IoT networks through anchor-aligned target construction, and analyzes how reference-classifier error, view construction, and sharpening affect the target.

Jin Yang, Xu-Feng Liu, Yong-Cai Hu et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory

We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifier evaluates 't Hooft anomaly matching, superpotential R-charge consistency, central-charge matching, and a bounded chiral-ring proxy. A claim that passes receives a consistency cert...

Xingyang Yu · 0 citations

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment

This paper proposes TRACE, which simulates harmful SFT trajectories to produce progressively corrupted model states, and optimizes a safety patch simultaneously across these states, and improves the safety rate by up to 77 percentage points over the best baselines, while maintaining comparable utility to the undefended...

Changyue Li, Jiaming He, Youliang Yuan et al. · 0 citations
#artificial intelligence Preprint Jun 2026

AutoDojo: A Generative Benchmark for Evaluating Prompt Injection Defenses in LLM Agents

AutoDojo is introduced, a generative benchmark built on AgentDojo and AgentDyn that generates IPI adaptively for a given agent and defense, and it is demonstrated that standard static benchmarks often significantly overestimate defense efficacy.

Xinhang Ma, Taoran Li, Chao-Wei Xiao et al. · 5 citations

Large Language Models Hack Rewards, and Society

To study this phenomenon, SocioHack is introduced, a sandbox of 72 societal environments, and it is found that within these environments, reward hacking naturally emerges and leads to regulatory loophole discovery.

Wei Liu, Xinyi Mou, Hanqi Yan et al. · 3 citations
#artificial intelligence Preprint May 2026

MemPoison: Bypassing Selective Memory Mechanisms to Plant Backdoors in LLM Agents

MemPoison is proposed, a novel memory poisoning attack that bypasses selective memory mechanisms in LLM agents, where an attacker can inject triggerable backdoors into the agent's long-term memory through dialogue interactions, thereby misleading its subsequent responses.

Hongtao Wang, Sean Yang, Yu Chen et al. · 5 citations

The End of Trust: How Agentic AI Breaks Security Assumptions

The Infinite Impostor is introduced, an attack model in which an autonomous agent interposes itself between two parties who already trust each other, hijacking an existing relationship rather than building a new one from scratch, and the governance tensions that arise when platforms become the regulatory substrate of d...

Osama Zafar, Alexander Nemecek, Erman Ayday · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Who Owns This Agent? Tracing AI Agents Back to Their Owners

AI agents increasingly act autonomously in the world, yet harmful behavior cannot be reliably traced to the account that deployed the agent. This creates an accountability gap across both benign and malicious settings: misconfigured or hijacked agents may cause unintended harm, while malicious operators may deploy agen...

Ruben Chocron, Doron Jonathan Ben Chayim, Eyal Lenga et al. · 0 citations

Watermarking Should Be Treated as a Monitoring Primitive

It is shown that even zero-bit watermarking supports internal attribution under per-entity multi-key deployments without explicitly encoding identity, and external identification in selected text and image configurations is demonstrated.

Toluwani Aremu, Nils Lukas, Jie Zhang · 1 citation

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.