This paper comprehensively studies effective prompt injection attacks against 14 widely used open-source and three closed-source LLMs on five attack benchmarks and proposes a straightforward and effective hypnotism attack, showing that this attack causes aligned language models to generate objectionable behaviors.
Jia-Wen Wang, Pritha Gupta, E. Hüllermeier et al.· 9 citations
A detection framework that models this attack chain by combining a text detector, verifiers specific to each stage, explicit rule-based risk signals, user intent and action consistency analysis, and a logistic decision policy is proposed, which achieves a mean F1 score under the strict threshold setting policy.
DUALLM, a dual-method pipeline that integrates two approaches based on a Large Language Model (LLM) and a fine-tuned small language model, achieves 87.4% accuracy and an F1-score of 0.875, significantly outperforming prior solutions.
Xing-Yu Li, Jue-Fei Pu, Yifan Wu et al.· Network and Distributed Syst...· 1 citation
By leveraging the better privacy-utility trade-off of PrivaTree, this work is able to train decision trees with significantly better robustness against backdoor attacks compared to regular decision trees and with meaningful theoretical guarantees.
D. Vos, Jelle Vos, Tianyu Li et al.· 6 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Intrusion detectors for small Internet-of-Things (IoT) devices are usually compressed by pruning and judged by overall accuracy, but it is shown that this hides a severe class-level failure, find its cause, and give low-overhead prevention and repair.
Encrypted training relies on keeping server-side updates low-degree. This constraint traditionally excludes models whose weights inhabit a compact Lie group (notably variational quantum circuits, where every trainable weight is an $\mathrm{SU(2)}$ rotation). Expressed in Euler angles or discrete alphabets, these update...
Marcel Mordarski, N. Mani, Arshad Patel et al.· 0 citations
Large language models (LLMs) have demonstrated promising performance in log anomaly detection, yet how their adaptation strategies, architectures, and deployment configurations affect detection effectiveness remains insufficiently understood. To investigate these factors, we conduct a systematic empirical analysis acro...
Best-of-N jailbreaking spends a query budget on surface variation, scrambling and recasing a request until one draw lands. We ask what a budget buys when its variance is moved into a structural channel instead, holding the search identical across both arms so the encoding is the only difference. Against SAGE, the stron...
Haoyu Zhang, Hanwen Liu, Yang Chen et al.· 0 citations
This work presents AgentXploit, a two-role auditing system that separates repository-level attack-path discovery from runtime exploitation and introduces AgentXploit-Bench, containing 72 reproducible vulnerabilities across 12 open-source AI-agent systems and frameworks.
Wei-Da Liang, Shi Qiu, Zhun Wang et al.· 0 citations
This paper proposes a blockchain-backed agentic security framework designed to safeguard the complete software development lifecycle (SDLC) while also securing the agentic AI components responsible for monitoring it. The framework coordinates a set of specialised security agents, covering source integrity, dependency a...
DeformView is introduced, a wide-baseline MV dataset with pixel-level annotations of geometric inconsistencies and DEFECt3R is proposed, a lightweight learning-based classifier that uses cross-view feature relationships to localize geometric inconsistencies at the pixel level.
Xander Staelens, Albéric Loos, Bert Ramlot et al.· 0 citations
JevAdvBench is introduced, to the authors' knowledge the first adversarial benchmark for RLCD models, with 812 typed questions over 66 scenarios, and a black-box attack suite of 9,744 single-edit variants that each edit one part of a request, with billed input tokens confirming that the edit reached the model.