Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Preprint Oct 2026

MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs

Quantized large language models are increasingly deployed on edge devices for their low latency and energy efficiency. However, model quantization weakens alignment safeguards, leaving qLLMs (quantized large language models) highly vulnerable to jailbreak attacks. To address this challenge, we present MOMAT (Mixture of...

Bo-Yang Li, Bingyu Shen, Wei-Hao Hong et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution

An autonomous cyber defender trained with reinforcement learning (RL) is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it does not eliminate th...

Harshith Doppalapudi, Nathaniel D. Bastian, Ankit Shah · 0 citations
#artificial intelligence Preprint Sep 2026

No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents

Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks. Reinforcement learning (RL) and large language models (LLMs) offer complementary mechanisms for the planning and execution such agents require, and prior work has combined them in hybrid...

A. Shaikh, Arunesh Sinha, Nathaniel D. Bastian et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents

Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an accumulated representa...

Hao-Yu Wang, Wei Zhao, Ye-Di Zhang et al. · 0 citations
#artificial intelligence Review Sep 2026

Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems

Multi-agent systems derive their capabilities from sharing evidence, delegating tasks, and combining information across agents. The same process creates a safety problem: contributions that are admissible in isolation can jointly enable a prohibited use. Blocking every sensitive action avoids disclosure but defeats the...

Yun-Bei Zhang, Saiyue Lyu, Janet Wang et al. · 0 citations
#artificial intelligence Review Sep 2026

Proof-Gated Signing: Solver-Checked Transaction Guards that Hold Under State Drift for Onchain AI Agents

AI agents that control wallets read attacker-reachable content, so they can be steered into proposing harmful transactions. The usual last line of defense is a pre-signing check: a static allowlist, an LLM reviewer, or a transaction simulation. All three share a gap: the check describes the chain state at check time, b...

Bravish Ghosh · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Authorization for Self-Modifying AI Agent Populations: Conserving Authority across Replacement, Forking, and Rollback

Self-modifying AI agents can replace, fork, and roll back identity-bearing software while descendants remain executable. Per-successor authorization does not constrain the resulting population: siblings may duplicate quotas, combine permissions, survive ancestor cuts, or overlap predecessors during promotion. We define...

Genliang Zhu, Chu Wang · 0 citations
#artificial intelligence Preprint Sep 2026

Actions with Receipts: Jointly Binding Claims, Evidence, and Execution for Replayable Tool-Agent Auditing

Tool-using agents can expose citations and execution logs while leaving a critical association unaudited: whether the claim shown to a user is the claim emitted by the committed execution and supported by the cited source. A valid citation and a valid trace can therefore remain individually well formed while being tran...

Miao-Bo Hu, Shu-Hao Hu, Xiao-Bo Guo et al. · 0 citations
#artificial intelligence Open access Sep 2026

Tokenized Key-Gated Adapter Routing: A Secure Access Control Mechanism Against Private Data Leakage in LLMs

Large language models (LLMs) are increasingly deployed in privacy-critical domains (e.g., healthcare, finance, and government), but their propensity to memorize and disclose personally identifiable information (PII) poses serious security and compliance risks. Existing defenses typically force a trade-off between model...

M. Shaaban, Mohamed Elmahallawy · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Multi-Jurisdictional Legal Identity Assurance for Capability Gating: A Design-Science Proposal for Tiered, Reusable Identity Assurance of Natural, Juridical, and Machine Entities

Identity assurance is the cost a digital system pays for dishonesty and uncertainty: it exists to make acts attributable when not everyone can be trusted at their word. A common way to pay that cost is flat maximum verification, asking each participant to meet a single high level of identification at entry, before any...

Walter Kurz · 0 citations
#artificial intelligence Preprint Sep 2026

The Cognitive Continuity Test: Verifying Governed State Transitions in Persistent AI Agents

Persistent AI agents revise beliefs, consolidate memory, and replace execution substrates. Similar successor states can accompany differently authorized transition claims, while legitimate development can change state substantially. We introduce the Cognitive Continuity Test (CCT), a policy-relative contract for verify...

Jun-Fei He, De-Ying Yu · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Verifier Can Leak the Answer: Diagnosability Before Optimization in Closed-Loop Agent Debugging

Agent developers increasingly compare prompts, tools, policies, and diagnosis algorithms through simulator-grounded verifiers. A verifier can nevertheless make a solver comparison vacuous: if its probes or predicates encode the target identity, an exact optimizer may appear effective without resolving any genuine ambig...

Peiying Zhu, Sidi Chang · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.