Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign alternatives on the...
XiaoYu Xu, Zi Liang, Min-Xin Du et al.· 0 citations
A taxonomy of agentic commerce fraud that separates five observation levels (agent reasoning, wire, settlement rail, counterparty, principal) from the request-level and history-level evidence available at each, and records which levels can observe which attacks.
Tool-augmented data agents rely on tool outputs for analytical decisions. Yet successful execution can return plausible but incorrect evidence, requiring agents to decide whether to trust or verify it. Understanding this failure requires examining both the evidence obtained through checking and the answer ultimately ad...
The EU AI Act introduces extensive compliance requirements for organizations that develop, deploy, or integrate AI systems. Many of these requirements are directly relevant to security and privacy, while also addressing closely related issues such as data governance, transparency, accuracy, and robustness. However, sta...
Zhen Tao, Alize Kahraman, Shidong Pan et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
The misaligned AI behaviors that led to the OpenAI-Hugging Face incident are identified and it is shown that an auditing agent can elicit similar behaviors given high-level qualitative descriptions and that RL is a promising direction to do so.
Stewart Slocum, Malayandi Palan, Christopher G. Chute et al.· 0 citations
DECEIVE (Deepfake Exploitation of Cognitive Engagement and Implicit Visual Evaluation), a framework that models how deepfake videos are validated as adversarial payloads through behavioral and neuro-physiological screening of viewers, and how attacks can be refined by selecting payloads that evade detection.
Cagri Arisoy, Mahmuda Huq, Amy W. Hays et al.· 0 citations
This paper presents AuditBench, a new benchmark dataset for evaluating the capabilities of LLMs at investigating security-related system audit logs. We design and use this benchmark to explore the performance of LLMs on four log-investigation tasks that incident response teams commonly perform, ranging from triaging al...
Aniket Anand, Yiwei Hou, Daniel Fields et al.· 0 citations
D-STT, a simple yet effective defense algorithm that identifies and explicitly decodes safety trigger tokens of the given safety-aligned LLM to activate the model's learned safety patterns, effectively preserves model usability by introducing minimum intervention in the decoding process.
Hao-Ran Gu, Han-Ding Wang, Yi Mei et al.· 0 citations
A per-token logits mask is applied that deterministically eliminates single-query column-use violations on the supported SQL fragment in a single decoding pass and achieves 0% Leakage Rate and Coverage up to 88.7% on Spider-CU, while staying within +10% tokens of direct prompting.
As Large Language Models (LLMs) are increasingly deployed in cross-linguistic contexts, ensuring safety across diverse regulatory and cultural environments has become a critical challenge. However, existing multilingual benchmarks largely rely on general risk taxonomies and machine-translated data, limiting evaluation...
Yunhan Zhao, Zhaorun Chen, Xingjun Ma et al.· 0 citations
Anyone with a public footprint leaks facts that were never stated, and language models make the inference cheap. We present a framework for measuring and reducing this inference exposure that runs on the owner's own CPU with no language model at analysis time, instantiated on organisations and on individuals. It starts...
When a language-model audit finds no leak, what is needed to certify non-leakage? We study guarantees over a declared prompt domain under an executable leakage criterion and decoding rule. For general bounded polynomial-time evaluators, a supplied leaking execution is polynomial-time checkable, while leak existence is...