Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Preprint Sep 2026

Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks

Thompson's"Reflections on Trusting Trust"showed that a compiler can be poisoned to reinsert its own backdoor, so that even recompiling clean source reproduces the Trojan. Today, substantial coding work is done by AI coding agents -- and increasingly, those agents generate new versions of themselves. We reconsider Thomp...

Franziska Roesner, Tadayoshi Kohno · 0 citations
#artificial intelligence Preprint Sep 2026

Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery

How does a multi-agent system evolve from a local deviation into collective loss of control? We propose an epidemic explanation organized around accidental mutation, contagion, and recovery. A spontaneous deviation creates a seed; communication enables other agents to adopt and retransmit its unsafe strategy; collectiv...

Xiang-Fan Wu, Zong-Hao Ying, Hui-Yu Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI

The model is given a formal basis by transplanting the beta-factor model of common-cause failure from reliability engineering, a seven-step protocol whose outputs a third party can verify, a structural detectability analysis of a procurement-controls agent audited at three grades, and a Monte Carlo study of the model.

Mohamed Chahine Ghanem · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Physics-Constrained Digital Twins for Sensor Integrity in Urban Pedestrian Flow: Detecting Stealthy False Data Injection with Conformal Guarantees

City pedestrian counting systems now feed economic indicators, planning decisions and safety operations, yet the twins built on top of them treat the incoming stream as ground truth. We study what happens when it is not. We formalise stealthy false data injection for city-scale pedestrian sensing, where the map from la...

Oscar Mogollon Gutierrez, Fatemeh Ghasemi, Mohammadhossein Homaei et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Making AI-Assisted Claims Independently Challengeable: Publication Authority and a Protocol for Falsifiable Publication Records

AI-assisted claims can appear authoritative when evidence, analysis, human authorization, presentation, and correction history refer to different states. Provenance, attestation, and transparency expose history but alone do not specify the publication transition examined here. We develop Publication Authority as an exa...

Torsten Tiltack, Yi-Fei Dong, Kun Yu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

What Breaks Local Watermarks? A Robustness Benchmark for Local Invisible Image Watermarking

This work presents the first systematic robustness benchmark for local watermark robustness across 55 image transformations, and finds that signal distortions are often tolerated by the strongest methods, while geometric misalignment and generative local edits can completely impair payload recovery.

Kai Yao, Bence Szilágyi, Sebestyén Kamp et al. · 0 citations
#artificial intelligence Preprint Sep 2026

The MAL Simulator: Cyber Operations Simulation based on Attack&Defense Graphs

The MAL Simulator, a cyber operation simulator based on the Meta Attack Language (MAL), found that the trained attacker policy could reach the designated targets more efficiently than the compared search methods, and that the trained defender agent induced lower costs than a naive heuristic agent under noisy alert cond...

Jakob Nyberg, Sandor Berglund, Andrei Buhaiu et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

A Cyber Range Evaluation of Autonomous Network Incident Response Agents

We test the performance of agents for automated network intrusion response in a cyber range intended for human operator training. The range implements an emulated networking environment with a variable network topology, red-team emulation and simulated user agents. The goal of the defensive agents is to prevent hosts i...

Jakob Nyberg, Teodor Sommestad, Andrei Buhaiu et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

SemanticAdv: Generating Adversarial Examples via Attribute-conditional Image Editing

Deep neural networks (DNNs) have achieved great success in various applications due to their strong expressive power. However, recent studies have shown that DNNs are vulnerable to adversarial examples which are manipulated instances targeting to mislead DNNs to make incorrect predictions. Currently, most such adversar...

Haonan Qiu, Chaowei Xiao, Lei Yang et al. · 0 citations
#cybersecurity Preprint Open access Sep 2026

Using Codebooks to Detect Cybercrime Topics in Text Narratives

In the United States, management of cybercrime-related consumer complaints increasingly falls on state and city governments given de-staffing of federal agencies. AI, and in particular, large language models (LLMs), shows promise for detecting cybercrime in text complaints, but often via specialized models that local g...

Shufan Chai, Liangliang Sun, Jessica Staddon · 0 citations
#artificial intelligence Preprint Sep 2026

RAG-CT: Mitigating Privacy Risks on Retrieval-Augmented Generation Systems via Scanning Prompt Distribution

Retrieval-Augmented Generation (RAG) has emerged as a powerful paradigm for improving the quality of generated contents of Large Language Models (LLMs) by grounding responses in external knowledge, thus reducing hallucinations and factual errors. However, recent studies have highlighted a critical vulnerability: advers...

Xingyu Lyu, Jia-Yi Wang, Jian-Feng He et al. · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.