Frontier AI models are rapidly gaining the ability to exploit vulnerabilities in complex pieces of software. The risk is not theoretical, as evidenced by recent sandbox escapes performed by frontier models at OpenAI and Anthropic. Discussions of how to sandbox inference stack components often focus on components other...
Sarah Radway, Andrew Cheng, Vijay Janapa Reddi et al.· 0 citations
While multimodal large language models (MLLMs) enable a wide range of image-text reasoning tasks, recent incidents indicate that they are vulnerable to illicit deployment and unauthorized distillation. Existing solutions for model provenance are typically confounded by shared language backbones in MLLMs and struggle to...
The Internet of Agents is expected to enable large numbers of autonomous agents to discover, verify, and collaborate with each other across heterogeneous platforms. However, current agent protocols mainly address tool invocation and inter-agent communication, leaving scalable agent registration, trustworthy identificat...
Song Zhang, Jiankang Yao, Hongtao Li et al.· 0 citations
As agent systems become more widely used, multiple agent sessions increasingly run alongside pre-existing user tasks in the same environment, sharing resources with limited capacity or mutually exclusive states. This creates a safety risk: when granted sufficient privileges, an agent may resolve a resource conflict by...
Yuejin Xie, Yu Li, Dadi Guo et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
The benchmark, failure-preserving contract, incident provenance, and governance controls needed to prevent infrastructure behavior from being misreported as model behavior and the benchmark, failure-preserving contract, incident provenance, and governance controls needed to prevent infrastructure behavior from being mi...
Hang Xiao, Chu-Hong Xu, Kai-Nan Zhou et al.· 0 citations
FARSIGHT (Financial Agent Robustness and Security Investigation and Global Holistic Testing), a framework that performs scheme-level evaluation of financial LLM agents on two axes: robustness under market turbulence and security against three attack types: attacks on information sources, attacks on agents, and agent-as...
The red-teaming methodology and highlighting new attack vectors aim to help defenders evaluate their mitigations against the possibility of persistent malign coding agents and prevent multi-context attacks at an acceptable cost.
Alex Remedios, Simon Storf, Fabien Roger et al.· 0 citations
AUDITPLAN is a single-model plan-then-answer approach where the model first emits a compact structured safety plan and then answers conditioned on it, enabling machine-checkable auditing while remaining hidden from users at deployment.
Sai Sri Pushpa Jampani, Kshitij Mishra, A. Ekbal· 0 citations
Large language models fine-tuned for network intrusion detection emit single-point predictions without statistical validity guarantees. Conformal prediction supplies a finite-sample coverage guarantee, but a threshold calibrated on clean traffic fails once an adversary perturbs controllable network features. We demonst...
This work presents PAPC, a platform-mediated mechanism that intercepts information-moving events before they update shared state or external channels and positions event-level mediation as a platform-governance primitive for agent-mediated online work.
Tao Huang, Guo-Xin Wu, Chen Hou et al.· 0 citations
Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical deployments. Existing external guardrail models remain blind to the model's internal workings, c...
Tool-augmented large language model (LLM) agents fail in a way no tool-selection or tool-security method addresses: they call tools that do not exist and pass arguments no schema declares. Existing defenses either pick the right tool (selection) or constrain what an agent may do with real tools (gating), both of which...