Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing...
CO Tiffany, Wen Zhang, E. Bagdasarian et al.· 0 citations
Several open research questions are identified, including how to efficiently record incidents and how to determine whether vulnerabilities and incidents generalize, and privacy requirements are summarized and research directions for the secure and trustworthy deployment of AI agents are outlined.
Anastasia Pustozerova, E. Bagdasarian, Luca Beurer-Kellner et al.· 0 citations
It is proved that, under an intuitive and experimentally supported assumption called distribution consistency, obfuscation can nullify the robustness of N-gram-based watermarks and motivate more semantics-aware alternatives.
Gehao Zhang, Eugene Bagdasarian, Juan Zhai et al.· arXiv.org· 0 citations
PiSAs (Privacy in Shared Agentic systems), a benchmark for assessing unintentional leaks with dual CI annotations, enables direct measurement of cross-user spillage across agentic system components and interfaces, such as outputs, inter-agent communication, and memory.
Shubham Gupta, N. Sepahvand, Abhinav Kumar et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.