Skip to content

Category

cybersecurity

1,032 papers

#artificial intelligence Preprint Open access Oct 2026

MiniScope: Authorizing Agents with Least-Privilege Permissions

AI agents are increasingly granted autonomous access to sensitive user data and third-party services, making effective permission management a critical security challenge. Existing permission models, however, typically rely on flat permission structures that fail to balance security with usability: fine-grained confirm...

Jinhao Zhu, Xiao Huang, Kevin Tseng et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Semantic Behavioral Watermarking: Paraphrase-Robust and Forgery-Resistant Provenance for LLM Agents

Behavioral watermarking embeds an owner identifier in an LLM agent's high-level action choices, giving provenance without touching output tokens. Prior agent watermarks break in two ways. First, all three prior schemes bind the watermark to the exact action symbol, so renaming a tool desynchronizes decoding even when t...

Suxin Ji, Hungtao Wan, Shaoxuan Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

When Tools Lie: Reliability of Mathematical Agents Under Corrupted Tool Feedback

Mathematical problem solving often requires deterministic computational steps that agents delegate to tools and implicitly trust. Yet tools can fail silently, returning plausible but incorrect results. How well can agents detect and correct corrupted tool call outputs? We study this through a controlled corruption fram...

Kavienan Jegatheesan, Gayathri Lihinikaduarachchi · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SkillPoison: Progressive Skill Poisoning via Successful Experiences

Self-improving LLM agents increasingly distill successful experiences into persistent, reusable skills. Existing skill attack methods corrupt this learning pipeline by injecting malicious triggers, behaviors, or false facts into individual experiences or extracted skills. However, such attacks are easily detected, and...

Lizhi Zhang, Xin He, Dianxuan Fu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?

Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch generation or vulnerability detection a...

Hang He, Li Wang, Hao Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Can Power Draw Constrain Covert Compute? Limits of Analogue Verification for AI Governance

Frontier AI treaties or agreements on limiting computation require external verification; an external auditor must be able to confirm how much computation actually ran and that parties are adhering to the agreement. Analogue, off-chip measurements such as power draw provide an information channel for verification. It i...

Tom Kimpson, Mauricio Baker, Emlyn Graham · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Polar: LLM-Powered Synthesis of Real-World Cyber Evidence for Prioritization and Mitigation

Cyber threat analysis increasingly depends on evidence distributed across vendor advisories, vulnerability databases, and threat intelligence sources. Turning these fragmented observations into timely decisions requires models to connect technical severity with evolving exploitation evidence and available defensive act...

Luoxi Tang, Yuqiao Meng, Ankita Patra et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

TARE: Weigh a Never-Poisoned Twin Before Reading Backdoor-Defense Costs

Backdoor-defense leaderboards print a clean-accuracy drop and read it as removal cost. Measured on the poisoned victim alone, the drop cannot separate removal from what the defense does to any model, and inherits the victim's start, which for three of BackdoorBench's sixteen attacks is a configuration file: WaNet, BPP...

Ruizhi Xu, Wei Xu, Sibo Zhu · 0 citations
#artificial intelligence Preprint Open access Oct 2026

APEX: Active Protection at Execution Boundaries for LLM Agents

Indirect prompt injection (IPI) hides adversarial instructions in content that large language model (LLM) agents read at runtime. As agents compose heterogeneous capability units, including Tools, MCP servers, and Skills, the carriers of injection multiply, and defenses built to recognize attack patterns fall behind th...

Xinran Zheng, Xin Fan Guo, Zhiqiang Hao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Dynamical low-rank equilibrium computation for stochastic games between advanced persistent threats and moving target defense

Moving target defense (MTD) against advanced persistent threats (APTs) in industrial control systems (ICS) has well-established game-theoretic formulations, but their practical value hinges on equilibrium computation, which faces two gaps: full-rank value iteration is prohibitively expensive at industrial scale, and th...

Tian Zijian, Zhang He, Chen Xinjie et al. · 0 citations
#cybersecurity Preprint Open access Oct 2026

The Amplifier Effect: Human-Factor Risks of AI-Suggested Correlation and Auto-Propagation in Multi-Framework GRC Self-Assessment

Multi-framework Governance, Risk and Compliance (GRC) platforms increasingly automate the link between an organisation's self-assessment answer and the compliance obligations that answer is said to satisfy. Cross-framework control mapping, AI-suggested question correlation, and automatic propagation of answers and evid...

Nikolaos Kekatos, Michael Ioannou, Marina Korgiala-Karyda et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Where Rules End and Judges Begin: Measuring the Judgment Boundary in Multi-Agent Systems Security

LLM-based multi-agent systems (MAS) engage tools, share memory, and delegate tasks, often encountering adversarial content. Current defenses for MAS are typically evaluated in isolation, focusing on one attack type at a time, which can lead to costly and hard-to-audit outcomes. This study organizes defenses into five p...

Shaswata Mitra, Raj Patel, Subash Neupane et al. · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.