With LLM watermarking already being deployed commercially, practical applications increasingly require multibit watermarks that encode more complex payloads, such as user IDs or timestamps, into the generated text. In this work, we propose a fundamentally new approach for multibit watermarking: introducing binomial enc...
Thibaud Gloaguen, Robin Staab, Mark Vero et al.· 0 citations
FraudBench is a multimodal benchmark for detecting AI-generated fraudulent refund evidence and shows that current MLLMs often recognize real-damaged evidence but fail on many fake-damaged subsets, with fake-damage detection rates far below the 50\% baseline on most generator subsets.
Experiments show that turn-level boundary supervision improves intervention localization, while reinforcement learning further improves the safety--utility trade-off.
Xin-Jie Shen, Rong-Zhe Wei, Pei-Zhi Niu et al.· 0 citations
MOSAIC-Bench (Malicious Objectives Sequenced As Innocuous Compliance), a benchmark of 199 three-stage attack chains paired with deterministic exploit oracles on deployed software substrates that treats both exploit ground truth and downstream reviewer protocol as first-class evaluation axes, is introduced.
J. Steinberg, Oren Gal· arXiv.org· 4 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Sigil, a server-enforced watermarking framework for U-SFL, defines a secret watermark constraint in the server-visible activation space and embeds the watermark into client-side models by injecting a watermark gradient into the gradients returned during training.
Zhengchunmin Dai, Jia-Xiong Tang, Peng Sun et al.· 0 citations
Many high-level security requirements are about the allowed flow of information in programs and are difficult to make precise because they involve selective downgrading. Notions from epistemic logic have emerged as a good approach to policy semantics but a robust general framework remains elusive. A paper appearing in...
This paper proposes a blockchain-enabled layered architecture for regulatory agent collaboration, comprising an agent layer, an off-chain computation layer, and an on-chain anchoring layer that establishes a systematic foundation for trustworthy, resilient, and scalable regulatory mechanisms in large-scale agent ecosys...
Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e.,"distill") their own models on these traces. Exi...
Shidan Javaheri, Alexander Panfilov, O. Britton et al.· 0 citations
This work introduces SEABench, a benchmark for studying endogenous misalignment arising from agent self-evolution, with 48 longitudinal task sequences that span multiple evolution surfaces, task domains, and harm types in a rich personal-assistant environment and shows that qualitatively different safety behaviors emer...
Saswat Das, Parvati Viswanathan, Daniel Donnelly et al.· 0 citations
The compiler must read modules a physics-based solver cannot build without; the coding agent must not read that intellectual property. The harness does not ship that rule. We classified fifteen read routes against a container, permission rules and a sandbox. None of the three can tell which program is reading.
Large language model (LLM) assistants increasingly help users draft content for third-party recipients. During private drafting, the user or the model may introduce an item and later remove or replace it. The model may remove the item from the intended content but reveal it again when stating the edit. We call such sta...
LLM agents solve complex tasks by iteratively changing files, invoking local tools, and interacting with remote services, which modifies state across their local environment and remote services. Today, agents and users must manage these changes explicitly, whether reverting exploratory actions or recovering from errone...
Jinnan Guo, Hao Mark Chen, Kapil Vaswani et al.· 0 citations