Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Preprint Open access Sep 2026

Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape

Frontier AI models are rapidly gaining the ability to exploit vulnerabilities in complex pieces of software. The risk is not theoretical, as evidenced by recent sandbox escapes performed by frontier models at OpenAI and Anthropic. Discussions of how to sandbox inference stack components often focus on components other...

Sarah Radway, Andrew Cheng, Vijay Janapa Reddi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Fingerprinting Multimodal Large Language Models

While multimodal large language models (MLLMs) enable a wide range of image-text reasoning tasks, recent incidents indicate that they are vulnerable to illicit deployment and unauthorized distillation. Existing solutions for model provenance are typically confounded by shared language backbones in MLLMs and struggle to...

Chao Huang, Meng Tong, Kejiang Chen · 0 citations
#artificial intelligence Preprint Open access Sep 2026

A Scalable Trust Discovery Architecture for the Internet of Agents

The Internet of Agents is expected to enable large numbers of autonomous agents to discover, verify, and collaborate with each other across heterogeneous platforms. However, current agent protocols mainly address tool invocation and inter-agent communication, leaving scalable agent registration, trustworthy identificat...

Song Zhang, Jiankang Yao, Hongtao Li et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

ClashBench: Conflicts Leading Agents to Seize and Harm

As agent systems become more widely used, multiple agent sessions increasingly run alongside pre-existing user tasks in the same environment, sharing resources with limited capacity or mutually exclusive states. This creates a safety risk: when granted sufficient privileges, an agent may resolve a resource conflict by...

Yuejin Xie, Yu Li, Dadi Guo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Trust, but Validate the Instrument: Auditing AI-Generated RTL Verification Plans on Authored Security-Regression Proxies

The benchmark, failure-preserving contract, incident provenance, and governance controls needed to prevent infrastructure behavior from being misreported as model behavior and the benchmark, failure-preserving contract, incident provenance, and governance controls needed to prevent infrastructure behavior from being mi...

Hang Xiao, Chu-Hong Xu, Kai-Nan Zhou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes

FARSIGHT (Financial Agent Robustness and Security Investigation and Global Holistic Testing), a framework that performs scheme-level evaluation of financial LLM agents on two axes: robustness under market turbulence and security against three attack types: attacks on information sources, attacks on agents, and agent-as...

Meng-Xiao Wang, Nitesh Saxena · 0 citations
#artificial intelligence Preprint Sep 2026

AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment

AUDITPLAN is a single-model plan-then-answer approach where the model first emits a compact structured safety plan and then answers conditioned on it, enabling machine-checkable auditing while remaining hidden from users at deployment.

Sai Sri Pushpa Jampani, Kshitij Mishra, A. Ekbal · 0 citations
#artificial intelligence Preprint Jul 2026

Robust Conformal Intrusion Detection via Traffic-Aware Calibration and Attack-Orbit Invariance

Large language models fine-tuned for network intrusion detection emit single-point predictions without statistical validity guarantees. Conformal prediction supplies a finite-sample coverage guarantee, but a threshold calibrated on clean traffic fails once an adversary perturbs controllable network features. We demonst...

Zhenpeng Li · 0 citations
#artificial intelligence Preprint Sep 2026

PAPC: Platform Mediation for Privacy-Propagation Externalities in AI-Mediated Workflows

This work presents PAPC, a platform-mediated mechanism that intercepts information-moving events before they update shared state or external channels and positions event-level mediation as a platform-governance primitive for agent-mediated online work.

Tao Huang, Guo-Xin Wu, Chen Hou et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical deployments. Existing external guardrail models remain blind to the model's internal workings, c...

Alizishaan Khatri, Chiquita Prabhu, Omkar Neogi · 0 citations
#artificial intelligence Preprint Sep 2026

Closed-World Resolution Against Tool Hallucination in LLM Agents

Tool-augmented large language model (LLM) agents fail in a way no tool-selection or tool-security method addresses: they call tools that do not exist and pass arguments no schema declares. Existing defenses either pick the right tool (selection) or constrain what an agent may do with real tools (gating), both of which...

Lax Iyer · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.