Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Preprint Open access Oct 2026

Cocoon: A System Architecture for Differentially Private Training with Correlated Noises

Machine learning (ML) models memorize and leak training data, causing serious privacy issues to data owners. Training algorithms with differential privacy (DP) have been gaining attention as a solution. However, these algorithms add noise at each training iteration and degrade accuracy, limiting their real-world adopti...

Donghwan Kim, Xin Gu, Jinho Baek et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Mitigating Watermark Forgery in Generative Models via Randomized Key Selection

Watermarking enables GenAI providers to verify whether content was generated by their models. A watermark is a hidden signal in the content, whose presence can be detected using a secret watermark key. A core security threat are forgery attacks, where adversaries insert the provider's watermark into content \emph{not}...

Toluwani Aremu, Noor Hussein, Munachiso Nwadike et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Counterfactual Evidence Audits Predict LLM-Agent Susceptibility to Ranked Context

LLM agents increasingly decide from evidence assembled by upstream systems: retrievers choose documents, recommenders choose posts, and memory systems choose prior events. Existing evaluations usually hold this evidence fixed, missing failures in which individually ordinary items form a systematically one-sided context...

Rana Muhammad Usman · 0 citations
#artificial intelligence Preprint Oct 2026

Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks

Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether it also influences the benchmark's measur...

Neeraj Karamchandani, Piyush Nagasubramaniam, Xin-Hong Xie et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Certified Mechanistic Edits: Behavioral Guarantees for Skill Removal and Preservation

Mechanistic edits (ablations, weight edits, activation steering) are the standard tools for unlearning a harmful capability from a neural network while preserving useful ones. Current approaches validate their effects only by testing, which can never cover an entire continuous region of inputs. Prior work at the interp...

Md Sazid Uddin, Md. Khairul Alam Mazumder, M. F. Mridha · 0 citations
#artificial intelligence Preprint Open access Oct 2026

LiBRA: Detection-Aware Image Watermark Removal via Bidirectional Latent Optimization

Digital watermarking supports source attribution for AI-generated images, but its reliability depends on resistance to removal attacks. Some attacks attempt to remove watermarks by forcing the decoded watermark to differ from the original. However, this can produce an inverted watermark that remains detectable, causing...

Saibo Ye, Huajie Chen, Xin Guo et al. · 0 citations
#artificial intelligence Preprint Oct 2026

EvoRiskBench: An Evolving Benchmark for Runtime Security Risks in Workspace Agents

Workspace agents combine large language models with execution harnesses to perform stateful, multi-step tasks that access or modify external resources. Existing benchmarks leave gaps in executable coverage of their runtime security risks, while evolving model capabilities, harnesses, tools, and threats motivate benchma...

Shi Kuang, Xue-Mei Luo, Kun Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

The Fragility of Trigger-Tag Mechanisms for Misuse Detection in Open-Weight LLMs

Open-weight language models can be downloaded, modified, and deployed beyond their developers' control, limiting the effectiveness of centrally enforced safeguards. Recent work has therefore proposed \emph{trigger-tag} mechanisms that produce a detectable signal when a model is used under a target condition, such as ge...

Toluwani Aremu, Manit Baser, Mohan Gurusamy et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Securing Computer-Use Agents Against Branch Steering Attacks

Modern Computer Use Agents (CUAs) directly interact with graphical user interfaces and execute third-party web tools, exposing them to indirect prompt injection across every rendered page and tool response. While the Dual-LLM pattern is the primary system-level architecture offering formal security guarantees - using a...

Giulio Zingrillo, Hanna Foerster, Ilia Shumailov et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Beyond Predefined Sinks: Security-Aware Dependency Analysis for LLM Agents

Large language model (LLM)-based agents increasingly connect model-generated decisions to security-sensitive software capabilities such as command execution, filesystem access, network communication, browser control, and external tools. Existing analyses often use predefined sensitive operations as anchors, but operati...

Hang Cui · 0 citations
#artificial intelligence Preprint Open access Oct 2026

AgentTrap: Stateful Feedback Deception against Autonomous Penetration Testing Agents

Autonomous penetration testing agents conduct multi-step attacks by continuously adapting their plans and actions to target responses. As a common defense, honeypots can be deployed to divert these agents from real assets by presenting decoy services, while also supporting attack tracing and active counterattacks. Howe...

Yuelin Wang, Jiongchi Yu, Yanbang Sun · 0 citations
#artificial intelligence Preprint Oct 2026

Pincer: Resource Authorization for Agents using a Digital Twin

Coding agents have become increasingly long-horizon, autonomous, reliant on general-purpose shell and maintain their own persistent memory for self-improvement. While these capabilities have made the agents powerful, they have also made them harder to defend against external adversaries. Defenses that restrict this arc...

Mayank Rathee, Alexander Stepanov, Shalin Madabhavi et al. · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.