Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Preprint Sep 2026

MATE: Policy-Aware Security Auditing for Mobile Agents via Synthesis-Driven Trajectory Learning

Mobile agents powered by foundation models now automate complex, multi-step workflows on real devices, but their trajectories can violate app-specific security policies. Existing trajectory-level defenses rely on LLM prompting or rigid rules, and thus fail to support fine-grained, natural-language policies that general...

Chang-Yue Jiang, Jia-Yi Wang, Xin Wen et al. · 1 citation
#artificial intelligence Preprint Sep 2026

From Capability to Assurance in Autonomous Penetration-Testing Harnesses: A Framework and Reference Implementation

Research on large language model agents for penetration testing is evaluated almost entirely by capability: whether the agent captures a flag or reproduces a proof of concept. That metric suits a benchmark but is silent on the properties that decide whether an autonomous agent can be used in an authorized engagement: w...

Joas Antonio dos Santos Barbosa · 1 citation
#artificial intelligence Preprint Sep 2026

Fairly Compensated Distributed Information Retrieval and Augmentation for AI Agents

A fairly compensated protocol for distributed information retrieval and augmentation in autonomous agent networks that enables retrieval agents to securely evaluate and rank candidate documents without learning their plaintext contents, while ensuring that data providers are compensated only when valid information is s...

Yi-Xiang Yao, Pasha Barahimi, Srivatsan Ravi · 0 citations
#artificial intelligence Review Sep 2026

Zero-Trust Authorization and Discovery for Enterprise MCP

LLM agents translate natural-language context, which may include attacker-controlled text, into privileged tool calls, so authorization must remain effective even when an agent is prompt-injected or adversarially steered. The Model Context Protocol (MCP) has become a widely adopted interface for this boundary, yet its...

Huang-Jian Li, Yu-Wei Wang, Srinivasan Manoharan · 1 citation
#artificial intelligence Preprint Sep 2026

Structured Decomposition for Reliable LLM-Generated Access Control Policies

This paper presents an LLM-based system that translates natural-language access control policies (NLACPs) into executable Rego code for Open Policy Agent (OPA). It provides a modular, end-to-end pipeline for policy detection, component extraction, schema validation, linting, compilation, and automated test generation a...

Vatsal Gupta, Darshan Sreenivasamurthy · 0 citations
#artificial intelligence Preprint Sep 2026

Agents That Edit Documents: Measuring Agentic PDF Forgery Against a Non-Agentic Control

AI agents that carry a multi-step computer task through on their own became ordinary tools in the past year, and the same autonomy is available to anyone whose task is harmful. We ask what that means for a relying party -- an insurer, a lender, an auditor -- whose evidence is a filed PDF. AgentForge-Bench measures how...

Si-Miao Ren, Ankit Raj, Tommy Duong et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models

We evaluate the adversarial robustness of three frontier large language models (LLMs) developed by Anthropic, Opus 4.8, Fable 5 and its successor Fable 5.1, against four families of automated jailbreak attack across 7,826 harmful intents spanning a ten-category harm taxonomy. Using the HackAgent red-teaming framework,...

Nicola Franco · 0 citations
#natural language process... Preprint Open access Sep 2026

BAIT: Boundary-Guided Disclosure Escalation LLM Jailbreaking via Self-Conditioned Reasoning

In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that elicits malicious information through internal disclosure by target large language models (LLMs), instead of external feedback from judge LLMs. BAIT first asks the model to identify the protection boundary, then require...

Xuan Luo, Yue Wang, Geng Tu et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

BreakFun: Jailbreaking LLMs via Object Instantiation under Simulated Code Execution

Large Language Models (LLMs) are widely used because they process structures, syntax and code well, but this same ability also makes them paradoxically vulnerable. We introduce BreakFun, a jailbreak method that frames a harmful request as code-execution simulation. The prompt gives the model a benign Python class defin...

Amirkia Rafiei Oskooei, Mehmet S. Aktas · 0 citations
#artificial intelligence Preprint Sep 2026

LLMs as Linguistic Chameleons: Decoupling Semantics and Structure for Privacy-Preserving Communication

As Large Language Model (LLM) APIs become increasingly integrated into privacy-sensitive workflows, ensuring inference-time privacy without compromising task utility remains a major challenge. Existing approaches preserve most of the original semantic content to maintain downstream performance, but this also leaves exp...

Yu-Zhu Mao, Liang Zhao · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Single-Model Injection: A Threat Model and Defense Architecture for Prompt Injection in Multi-Agent Systems

Existing prompt injection research focuses on single-model chatbot scenarios, where an attacker manipulates one LLM through crafted input. Multi-agent systems amplify this threat through three mechanisms absent from single-model settings: inter-agent message passing creates injection channels invisible to perimeter def...

Rudrendu Kumar Paul, S. Nandy · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.