Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Preprint Aug 2026

A Formal Analysis of Agent Payment Protocols

This work formalizes four representative agent payment protocols: x402, MPP, ACP, and AP2 in Tamarin, and constructs source-grounded models that capture each protocol's roles, state, trust assumptions, and lifecycle transitions.

Ke Jiang, Mo-Han Yu, Yuan-Yi-Chun-Min-Chieh Chang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades

The theory that explains the blindness, the theory that explains the blindness, the two-population conservation law, under which every in-loop metric improves while true quality does not, and a synthetic study that validates the mechanism are given.

D. Rajput, Nirdesh Chauhan, S.Rao Kosaraju · 0 citations
#artificial intelligence Preprint Sep 2026

Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts

3R-Bench (Refusal, Repetition, and Revision), a benchmark of 150 real-world cybersecurity requests augmented with two adversarial conversational settings, is introduced and eight LLMs are evaluated, finding that prior assistant behavior strongly changes responses to an unchanged request.

Rui Yang, Yang Hong, Yi-Chao Xu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets

OpenAgentFlow is presented, a control-plane/action-plane architecture that establishes the action-commit boundary as a shared enforcement interface and shows that a shared action-commit boundary provides a practical basis for system-wide governance across heterogeneous agent execution paths.

Dong-Sheng Chen, Xiang-Yu Zhao, Xin Yao et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls

Gpt-oss-120b, a mixture-of-experts model with only $\sim$5.5B active parameters per token, carries the full state across all calls and returns the correct digest on a majority of completed runs, and localizes the residual failures by origin, separating state-carrying from arithmetic and from serving.

Dheeraj Mohandas Pai, Xian-Feng Lu · 0 citations

Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot

Experiments show that Knowledge Trap reduces surrogate Agreement by 6.2\% on average without degrading legitimate-user accuracy, outperforming existing defenses that impose measurable user impact, and suggest that defending knowledge-space traversal is a practical direction for mitigating LLM extraction attacks.

Yu-Yang Dai, Yushun Dong · 1 citation

Adversarial Trust Poisoning in Vehicular Collaborative Perception

This work presents TrustFlip, a novel attack that weaponizes consistency-based defenses to poison the trust assigned to benign vehicles, and introduces TrustReflect, a lightweight self-reflection mechanism that marks disputed regions as uncertain and excludes them from trust evaluation, reducing the attack success rate...

Yu-Tong Liu, Chenyi Wang, Ming F. Li et al. · 0 citations
#artificial intelligence Review May 2026

Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety

It is argued that evaluation should move beyond static question answering and model accuracy towards workflow-level assessment of evidence quality, tool use, policy and invariant compliance, staged execution, recovery, calibration, cost, and human intervention.

Muhammad Bilal, Jon Crowcroft, Rui-Zhi Wang et al. · 3 citations · ⚡1

Secret Stealing Attacks on Local LLM Fine-Tuning through Supply-Chain Model Code Backdoors

This work introduces a deterministic full-chain memorization mechanism that locks onto token-level secrets in dynamic computation flows via online tensor-rule matching, and leverages value-gradient decoupling to stealthily inject attack gradients, overcoming gradient drowning to force model memorization.

Zi Li, Tianyang Zhou, Wenze Li et al. · 0 citations

AgenTRIM: Tool Risk Mitigation for Agentic AI

AgenTRIM is introduced, a framework for detecting and mitigating tool-driven agency risks without altering an agent's internal reasoning that provides a practical, capability-preserving approach to safer tool use in LLM-based agents.

Roy Betser, Amit Giloni, Shamik Bose et al. · 14 citations · ⚡2

Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks

JAWS-BENCH(Jailbreaks Across WorkSpaces), a benchmark spanning three escalating workspace regimes mirroring attacker capability, is presented, indicating that JAWS-BENCH can be reused across multiple agent frameworks.

Shoumik Saha, Jifan Chen, Sam Mayers et al. · 8 citations · ⚡2

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.