Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Preprint Oct 2026

Out of Sync, Out of Sight: Phantom State Attacks against IIoT Intrusion Detection

Machine learning-based intrusion detection systems (IDS) are critical for securing Industrial Internet of Things (IIoT) environments. Most adversarial research against them perturbs the feature vector or the traffic that produces it, and depends on gradient access, repeated model queries, or a learned model of benign t...

Sabrine Ennaji, E. Benkhelifa, Nadia Kabachi · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SideKernel: A Usable microVM Sandbox for AI Coding Agents on macOS

AI coding agents are untrusted system components, yet they require autonomy on the developer machines they run on. This contradiction is a security problem. Sandboxes provide an isolated environment, but for local macOS development, the existing local, open-source options for AI coding agents are few in number and cumb...

Dimitrios Prasakis · 0 citations
#artificial intelligence Preprint Oct 2026

Mitigating Private Data Leakage in LLMs with Whiteout

Modern large language models (LLMs) are trained on massive, largely unfiltered datasets, including content scraped from nearly every accessible website and user inputs. As a result, LLMs often memorize and reproduce personally sensitive information (PSI) such as birth dates, phone numbers, and home addresses. This lead...

Anna Yoo Jeong Ha, Ronik Bhaskar, Hai-Tao Zheng et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Hop-Decayed Influence: New Vulnerabilities of Structural Auxiliary Indexing in GraphRAG Pipelines with LLM

GraphRAG pipelines construct auxiliary structures during offline indexing--semantic summaries, hierarchical edges, and pre-computed scores--that determine how retrieval is prioritised at query time. Prior attacks target only instance-level components (nodes, edges, triples), overlooking these schema-level structures. W...

Jisung Park, John Le, Heath Cooper · 0 citations
#artificial intelligence Preprint Oct 2026

MIRROR: Multipath Quorum Integrity for LLM Multi-Agent Communication

Inter-agent communication is central to Large Language Model Multi-Agent Systems (LLM-MAS), but it introduces an underexplored vulnerability: Agent-in-the-Middle (AiTM) attacks that manipulate messages in transit without compromising the agents themselves. Prior work reports Attack Success Rates (ASR) approaching 100%...

Ryuichi Yamafuji Lun, Jing-Zhen Wang, Shreyas Kolte et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Reasoning Models Are Accurate but Unsound on Identification

A reasoning model asked whether a causal effect is recoverable from observational data can fail in two ways: it refuses an identifiable query or answers a nonidentifiable one. The latter is more consequential, as no observational data can validate the claimed formula. Measuring this failure requires queries that are pr...

Arman Behnam, Binghui Wang · 0 citations
#artificial intelligence Preprint Oct 2026

Frequency Is Not Sensitivity Identifying Safety-Sensitive Experts in Sparse MoE LLM

Suppressing a small set of routed experts can weaken the safety behavior of a sparse Mixture-of-Experts (MoE) language model without retraining. Which experts to suppress is therefore a security question, and the usual answer is activation frequency, but frequency measures use, not influence. We test an alternative: ro...

Md Nurul Absar Siddiky, Liu-Wan Zhu, Yi Dong · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses

Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries an injection. System-1 decision models answer such questions in a single forward pass with class probabilities, promising large cost and latency savings over LL...

Jiawei Li · 0 citations
#natural language process... Preprint Oct 2026

OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents

LLM agents with tool-calling capabilities can access external services and private user data, but they may retrieve more information than a user's request explicitly requires. We study this behavior in structured tool-calling agents and term it proactive over-authorization. This setting differs from filesystem-level co...

Tao-Lin Zhang, Jiu-Heng Wan, Han-Yu Wang et al. · 0 citations

AuraForge: Scaling Security Supervision for Training Coding Agents

Coding agents are now proficient enough to generate complex software applications from a single prompt. As their capabilities have grown, human oversight has increasingly shifted from line-by-line code review toward hands-off evaluation of outcomes. However, recent studies have shown that such a transition exposes a cr...

Dan-Qing Wang, Song-Wen Zhao, Harsh Sharma et al. · 0 citations
#natural language process... Preprint Oct 2026

DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text

Robust detection of AI-generated text under deployment conditions is challenging: distribution shifts across domains and generators, adversarial perturbations of the input surface, and the absence of target-domain labels for threshold calibration all degrade detectors that perform well in-domain. We present DeBERTa-Con...

Mohamed Mady, Yu-Pei Li, Johannes Reschke et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective

Training generative machine learning models to produce synthetic tabular data has become a popular approach for enhancing privacy in data sharing. As this typically involves processing sensitive personal information, releasing either the trained model or generated synthetic datasets can still pose privacy risks. Yet, r...

Georgi Ganev, Emiliano De Cristofaro · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.