Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Preprint Open access Sep 2026

Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique

We propose a system for marking sensitive or copyrighted texts to detect their use in fine-tuning large language models under black-box access with statistical guarantees. Our method builds digital ``marks'' using invisible Unicode characters organized into (``cue'', ``reply'') pairs. During an audit, prompts containin...

Yanming Li (PETSCRAFT), C\'edric Eichler (PETSCRAFT), Nicolas Anciaux (PETSCRAFT) et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

FedNIA: Noise-Induced Activation Analysis for Mitigating Data Poisoning in Federated Learning

Federated learning systems are increasingly threatened by data poisoning attacks, where malicious clients compromise global models by contributing tampered updates. Existing defenses often rely on impractical assumptions, such as access to a central test dataset, or fail to generalize across diverse attack types, parti...

Ehsan Hallaji, Roozbeh Razavi-Far, Mehrdad Saif · 0 citations
#artificial intelligence Preprint Sep 2026

A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem

Agents using the Model Context Protocol (MCP) rely on semantic matching to select tools from third-party servers, exposing a semantic supply-chain risk through attacker-controlled metadata and outputs. We introduce A2M (Attraction-to-Manipulation), a two-stage black-box framework for hijacking MCP agents. The Attractio...

Lai-Zhen Li, Xuan Wang, Pei-Cheng Zhao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen

Large language models (LLMs) are increasingly applied to the automated repair of C/C++ security vulnerabilities, and compile rate is a commonly reported proxy for progress: whether the generated patch compiles. We argue that compile rate is a scientifically unreliable metric for single-function vulnerability repair, an...

Om Nepal, Sushant Aryal, Oluseyi Olukola et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Alignment to Access Control: A Framework for GenAI Policy Enforcement

Generative AI (GenAI) applications have flourished enabling users to chat with large language models, and to create agents to act on their behalf for a variety of tasks. The pace of development of capabilities in this field is incredibly fast with security and safety taking a back seat. Unfortunately, the slower pace a...

Nathalie Baracaldo · 0 citations
#artificial intelligence Preprint Sep 2026

Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models

It is found that Astra exhibits token-efficient directed reasoning, selecting a correct trajectory earlier, while resolving elementary steps internally and externalizing only crucial reasoning, which provides a behavioral lens on frontier-model reasoning beyond benchmark scores.

Xiao-Yu Luo, Tao Ren, Wen-Rui Yu et al. · 0 citations
#artificial intelligence Review Sep 2026

On the security and privacy of LLMs in Mobility

A clear gap is identified between strong optimization performance and regulatory adherence in AI Act compliance, suggesting compliance is limited less by technology than by a focus on static performance over lifecycle safety, and underscoring an urgent need for security-by-design in safety-critical intelligent transpor...

Mauro Conti, Lorenzo Perinello, Umberto Salviati · 0 citations
#artificial intelligence Conference Open access Mar 2026

Dynamic Deep Prompt Optimization for Defending Against Jailbreak Attacks on LLMs

Experiments demonstrate that DDPO significantly outperforms static prompt optimization methods, particularly on weakly aligned models and when handling semantically ambiguous benign prompts, successfully distinguishing them from genuinely harmful requests.

Doniyorkhon Obidov, H. Yu, Xiaolong Guo et al. · 4 citations
#artificial intelligence Preprint Open access Sep 2026

Silent Sabotage: Internal State Triggered Backdoor Attacks on LLM-Powered Robotic Systems

It is demonstrated that an attacker can embed a stealthy backdoor into an LLM-based robot controller by manipulating its instructions, triggered not by an external cue, but by a specific, rare sequence of the robot's own past actions.

Doniyorkhon Obidov, Shivayogi Akki, Cheng-Qiu Tan et al. · 2 citations
#artificial intelligence Preprint Aug 2026

Refusing Everything Looks Safe: Restoring the Benign Arm to Encoded-Prompt Evaluation

Encoded-prompt attacks are evaluated almost entirely on their harmful arm: a benchmark sends obfuscated harmful requests and reports how often the model complied, and shows that on two of the four models the loss is caused by the protocol rather than by the character transformation, and on a third by the characters.

Haoyu Zhang, Hao-Wen Xu, Xiao-Mao Luo et al. · 0 citations
#artificial intelligence Preprint Aug 2026

The Uncontrolled Variable: Vision-Language Refusal Is Conditioned on the Image-Attachment Interface, and Not Robust to Irrelevant Image Properties

It is shown that aligned vision-language models also condition refusal on a property of a request's form: whether an image is attached, holding everything the request asks fixed, which shifts benign refusal by tens of points.

Haoyu Zhang, Yi Feng, Han-Wen Liu et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

StepTrigger: Contact-State-Triggered Backdoor Attacks on VLM-Powered Legged Robots

Large language models and vision-language models are increasingly used as high-level planners in robotic systems, using task goals and sensor summaries to select navigation or manipulation actions. This creates a new backdoor surface: a compromised planner can behave normally in most runs, yet change its target selecti...

Jiageng Zhang, Doniyorkhon Obidov, Kaichen Yang · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.