Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Preprint Open access Oct 2026

AI Security Research Should Better Incentivize Defense Research

This work examines an imbalance in artificial intelligence (AI) security research: the field tends to produce more work on attacking AI systems than on defending them. Drawing on related academic papers, we find biased attack-to-defense ratios across subfields, including federated learning, speech recognition, membersh...

Youqian Zhang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs

Inference optimization aims to minimize the latency and resource consumption of LLM inference while preserving output quality, making large-scale deployment practical and cost-effective. However, optimized execution can introduce small numerical inconsistencies from the original model. We reveal that this inconsistency...

Yifei Wang, Yida Yang, Tianlin Li et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Five Production LLMs

Multi-agent LLM systems now read documents, web pages and tool results on behalf of users, yet their resistance to prompt injection is usually reported as one number: did the attack succeed? We introduce a kill-chain canary method that plants a unique token in every injected payload and records the furthest of four sta...

Haochuan Kevin Wang, Zechen Zhang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

GraphToxin: Reconstructing Full Unlearned Graphs from Graph Unlearning

Graph unlearning (GU) has emerged as a promising solution to comply with "the right to be forgotten" regulations by enabling the removal of sensitive information upon request. However, this solution is not foolproof. The involvement of multiple parties creates new attack surfaces, and residual traces of deleted data ca...

Ying Song, Balaji Palanisamy · 0 citations
#artificial intelligence Preprint Sep 2026

CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion

Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode - safety generalization lag - where alignment trained predominantly on natural language fails to trans...

Zheng Liang, Hai Huang, Wen-Tao Chen · 0 citations
#artificial intelligence Preprint Sep 2026

Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents

Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skill...

Tobias Kaisar, Aritra Dhar · 0 citations
#artificial intelligence Preprint Sep 2026

Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks

As Large Language Model (LLM) agents are increasingly deployed in complex environments, multi-turn interaction attacks have become a significant security challenge. Existing detection methods typically rely on historical context. However, this retrospective logic struggles to identify deep malicious intents that are sp...

Ze-Zhong Wang, Xue-Yang Tang, Rui Lian et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Learning Normal Diffusion Dynamics for Backdoor Defense in Text-to-Image Models

Backdoor attacks pose a serious threat to the secure deployment of text-to-image (T2I) diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal representations, which may limit their generalizability with the emergence of increasingly diverse attack mechanisms. In this...

Jun-Jian Li, Xiao-Long Liu, Peng Sun et al. · 1 citation
#artificial intelligence Preprint Open access Oct 2026

ActionGuard: Tool Call Authorization under Poisoned Skills

LLM-based agents extend their capabilities through third-party skills that provide task-specific instructions, scripts, and tool-use procedures. However, malicious instructions inserted into an otherwise benign skill can cause a benign user request to trigger dangerous Tool Calls, including data exfiltration, file dele...

Jihun Han, Yejin Jang, Byung Il Kwak et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Hiding in Plain Sight: Decoupling Pretext from Actuation for Skill Poisoning in LLM Agents

LLM agents increasingly rely on reusable Skills for complex, multi-step tasks, creating a critical supply-chain attack surface where poisoned Skill content steers agent decision loops under benign requests. Existing skill poisoning attacks either colocate actuation with its contextual pretext or distribute actuation ac...

Wenxin Wu, Lingyong Yan, Lei Sha et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Faithful Dual-constrained Erasure for Robust LLM Safety Alignment

Machine unlearning has emerged as a crucial mechanism for removing hazardous knowledge and enforcing safety alignment in Large Language Models (LLMs). However, recent studies reveal a persistent security risk: unlearned models remain highly vulnerable to retraining attacks, where suppressed malicious behaviors rapidly...

Jia-Qing Li, Shi-De Zhou, Zhi-Bo Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RAGScope: A Leakage-Controlled, Cost-Aware Evidence-Gating Protocol for RAG Hallucination Triage

Retrieval-augmented generation (RAG) systems need inexpensive ways to route generated answers: accept low-risk outputs, review uncertain ones, and reserve strong verifiers for the expensive tail. We present RAGScope, a leakage-controlled protocol for evaluating local evidence gates that use only the task input, retriev...

Zeming Liu, Qibai Chen, Jingtao Zhang et al. · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.