Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Review Oct 2026

Does AI Help Cyber Attackers or Defenders? Evidence from Nonpublic Vulnerabilities and Subsequent Attacks

The release decision for frontier AI systems increasingly relies on cyber capability benchmarks, yet public vulnerability benchmarks can expose agents to previously published advisories, exploits, and fixes, making it difficult to distinguish prior exposure from capability on unseen vulnerabilities. We evaluate open-we...

Tobias Heldt, Matthew Turk, Christoph R. Landolt et al. · 0 citations
#artificial intelligence Preprint Oct 2026

GCTAuto-encoder: A Cross modal Framework for Security Flaw Detection in IoT Networks

IoT encompasses diverse physical entities, from smart home devices to autonomous vehicles, creating a complex environment with heterogeneous security models. This heterogeneity makes IoT sub-systems vulnerable to various network attacks. Modern security systems must therefore be more robust to ensure security and priva...

Najmieh Sadat Safarabadi · 0 citations
#artificial intelligence Preprint Oct 2026

RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's o...

Mohamed Dhouib, Clément Elliker, Alexi Canesse et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Where Did the Repair First Go Wrong? Localizing the Origins of Silent Failures in Agentic Vulnerability Repair

Localizing where an LLM-based agent first fails to uphold security during a repair can show which stage of its workflow needs an additional safeguard. This is difficult for silent failures, which are patches that pass syntactic and functional checks but still contain a security vulnerability. Because such patches give...

Wenji Bai, Muhammad Waseem, Zeeshan Rasheed et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Cross-Lingual Transferability of Training Data Extraction Attacks to Recover Memorized PII

The robustness of Personally Identifiable Information (PII) protection in Large Language Models (LLMs) is a critical concern, yet the risks associated with cross-lingual data extraction remain under-explored. This study evaluates the vulnerability of English-centric and multilingual models to Training Data Extraction (...

Alexandru Nazare, Agnese Profico, Nicol\`o Vania et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Backdooring Sparse Autoencoders

Sparse autoencoders (SAEs) are increasingly used not only to interpret language models but also to intervene on their internal representations. We show that this creates a supply-chain attack surface: a maliciously modified SAE can induce attacker-chosen behavior when inserted into the forward pass of an otherwise unch...

E. Ahlers, Daniel Passon, Tobias Kiecker et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Runaway Reaction: When Benign Skills Compose into Malicious Behavior

Agent skills package task-specific knowledge and procedures that can be composed to support complex agent tasks, while public marketplaces provide a growing pool of reusable skills. Existing security vetting, however, largely evaluates skills in isolation, leaving composition-induced risks underexplored. Such risks ari...

Zunlong Zhou, Ziyuan Yang, Mengyu Sun et al. · 0 citations
#artificial intelligence Preprint Oct 2026

SimpleMark: Fast Multi-Bit Text Watermarking under f -Divergence Constraints

We introduce a framework for multi-bit text watermarking with security defined directly through $f$-divergence from the base language model distribution. Unlike prior approaches that focus on average-key distortion-freeness or a particular statistical distance, our formulation supports general $f$-divergences, includin...

Benjamin D. Kim, Wan-Rong Zhang, Wei-Tong Ruan et al. · 0 citations
#artificial intelligence Preprint Oct 2026

AgentDoxx: Agentic Re-identification of Anonymized Text with Web Search

As Large Language Models (LLMs) gain tool use capabilities such as web search, they can retrieve and cross-reference public information, creating privacy risks beyond memorization. One manifestation is re-identification: linking an anonymized interview transcript to a named individual. Yet without ground-truth identiti...

Jia-Ning Wen, Tian-Shi Li · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Don't Judge an LLM Only by Its Activations: Discovering Suppressed Safety Features via Counterfactual Activation Potential

Mechanistic interpretability has emerged as the primary means to understand safety behavior of LLMs. However, existing tools primarily focus on the activating neurons or features of a model. The role of the remaining large set of inactive components is invisible to such methods. This work demonstrates that the inactive...

Swadesh Swain, Sanghamitra Dutta · 0 citations
#artificial intelligence Open access Oct 2026

AutoDP-LLM: automating data pre-processing for intrusion detection systems using large language models

The increasing complexity and scale of modern cyber-attacks demand intelligent and computationally efficient Intrusion Detection Systems (IDS). However, designing effective data pre-processing pipelines traditionally involves substantial trial-and-error effort and repeated evaluation of alternative configurations. For...

Bao-Phong Nguyen, G. Pham, Thai-Duong Do et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Reflections and Fragments: Securing LLMs Against Sequential Mosaic Attacks

Self-play red-teaming improves language-model safety by pitting attacker and defender roles against each other in a zero-sum game. However, real adversaries increasingly use mosaic attacks: multi-turn sequences whose individual fragments are innocuous in isolation yet assemble into a harmful payload. We develop a theor...

Emanuele La Malfa, Saar Cohen, Gabriele La Malfa et al. · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.