We propose a system for marking sensitive or copyrighted texts to detect their use in fine-tuning large language models under black-box access with statistical guarantees. Our method builds digital ``marks'' using invisible Unicode characters organized into (``cue'', ``reply'') pairs. During an audit, prompts containin...
Yanming Li (PETSCRAFT), C\'edric Eichler (PETSCRAFT), Nicolas Anciaux (PETSCRAFT) et al.· 0 citations
Federated learning systems are increasingly threatened by data poisoning attacks, where malicious clients compromise global models by contributing tampered updates. Existing defenses often rely on impractical assumptions, such as access to a central test dataset, or fail to generalize across diverse attack types, parti...
Agents using the Model Context Protocol (MCP) rely on semantic matching to select tools from third-party servers, exposing a semantic supply-chain risk through attacker-controlled metadata and outputs. We introduce A2M (Attraction-to-Manipulation), a two-stage black-box framework for hijacking MCP agents. The Attractio...
Lai-Zhen Li, Xuan Wang, Pei-Cheng Zhao et al.· 0 citations
Large language models (LLMs) are increasingly applied to the automated repair of C/C++ security vulnerabilities, and compile rate is a commonly reported proxy for progress: whether the generated patch compiles. We argue that compile rate is a scientifically unreliable metric for single-function vulnerability repair, an...
Om Nepal, Sushant Aryal, Oluseyi Olukola et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Generative AI (GenAI) applications have flourished enabling users to chat with large language models, and to create agents to act on their behalf for a variety of tasks. The pace of development of capabilities in this field is incredibly fast with security and safety taking a back seat. Unfortunately, the slower pace a...
It is found that Astra exhibits token-efficient directed reasoning, selecting a correct trajectory earlier, while resolving elementary steps internally and externalizing only crucial reasoning, which provides a behavioral lens on frontier-model reasoning beyond benchmark scores.
Xiao-Yu Luo, Tao Ren, Wen-Rui Yu et al.· 0 citations
A clear gap is identified between strong optimization performance and regulatory adherence in AI Act compliance, suggesting compliance is limited less by technology than by a focus on static performance over lifecycle safety, and underscoring an urgent need for security-by-design in safety-critical intelligent transpor...
Mauro Conti, Lorenzo Perinello, Umberto Salviati· 0 citations
Experiments demonstrate that DDPO significantly outperforms static prompt optimization methods, particularly on weakly aligned models and when handling semantically ambiguous benign prompts, successfully distinguishing them from genuinely harmful requests.
Doniyorkhon Obidov, H. Yu, Xiaolong Guo et al.· AAAI Conference on Artificia...· 4 citations
It is demonstrated that an attacker can embed a stealthy backdoor into an LLM-based robot controller by manipulating its instructions, triggered not by an external cue, but by a specific, rare sequence of the robot's own past actions.
Doniyorkhon Obidov, Shivayogi Akki, Cheng-Qiu Tan et al.· 2 citations
Encoded-prompt attacks are evaluated almost entirely on their harmful arm: a benchmark sends obfuscated harmful requests and reports how often the model complied, and shows that on two of the four models the loss is caused by the protocol rather than by the character transformation, and on a third by the characters.
Haoyu Zhang, Hao-Wen Xu, Xiao-Mao Luo et al.· 0 citations
It is shown that aligned vision-language models also condition refusal on a property of a request's form: whether an image is attached, holding everything the request asks fixed, which shifts benign refusal by tens of points.
Haoyu Zhang, Yi Feng, Han-Wen Liu et al.· 0 citations
Large language models and vision-language models are increasingly used as high-level planners in robotic systems, using task goals and sensor summaries to select navigation or manipulation actions. This creates a new backdoor surface: a compromised planner can behave normally in most runs, yet change its target selecti...