It is argued that for applications processing sensitive user data (medical, legal, journalistic, or personal), the Web-CLI should be the default architecture, as it makes data locality an independently verifiable technical property rather than a policy promise.
A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent is held fixed and only its meaning-preserving surface form varies, does the canonical-form score estimate model behavior well, and how much...
Yongxin Zhou, Jun-Wei Yao, Yuanzhe Liu et al.· 1 citation
This work forms model extraction monitoring as benign-calibrated traffic-window distribution testing: embed incoming queries into a semantic space and test whether their aggregate distribution deviates from historical benign traffic.
A novel function hijacking attack (FHA) that manipulates the tool selection process of agentic models to force the invocation of an attacker-chosen function and is largely agnostic to the context semantics and remains effective across domains and function sets.
A trajectory-aware evolutionary search method, T-MAP, which leverages execution trajectories to guide the discovery of adversarial prompts and enables the automatic generation of attacks that not only bypass safety guardrails but also reliably realize harmful objectives through actual tool interactions.
Hyomin Lee, Sangwoo Park, Yumin Choi et al.· arXiv.org· 8 citations· ⚡1
This work proposes a general framework that utilizes weaker single-layer watermarks to preserve the entropy required for effective multi-layer ensembling and demonstrates that this counter-intuitive strategy mitigates signal decay and consistently outperforms strong baselines in both detectability and robustness.
Ruibo Chen, Yihan Wu, Xuehao Cui et al.· arXiv.org· 2 citations
This work introduces a novel approach, \textit{CanaryTrace}, to safeguard the ownership of text datasets and effectively detect unauthorized use by RA-LLMs, and demonstrates high query efficiency, detectability, and consistency, along with minimal perturbation to the original dataset, all without compromising the perfo...
Yepeng Liu, Xuandong Zhao, D. Song et al.· arXiv.org· 15 citations· ⚡1
This work introduces a clinically grounded framework that evaluates leakage along a graded axis of adversarial access, ranging from publicly inferable demographics to leaked note fragments, and provides a practical, reusable framework for contextual privacy evaluation of medical LMs.
Sasha Ronaghi, Sana Tonekaboni, Lena Stempfle et al.· arXiv.org· 0 citations
This work introduces AgentREVEAL, a diagnostic framework for analyzing retrieval-induced safety degradation in LLM agents, and uncovers the Safe Source Paradox, a safety-utility trade-off for retrieval-enabled agents.
Aditya Nawal, Manit Baser, M. Gurusamy· arXiv.org· 0 citations
This work finds that jailbreak robustness is highly fragile to operational-state variation: even when the attack remains fixed, changing only an ordinary system prompt not designed to affect safety can dramatically alter attack success rates.
Yuna Park, Hwang Youn Kim, Yujin Kim et al.· 0 citations
CIPR (Coding In Poisoned Repos), the first benchmark that systematically varies PLCs in poisoned real-world repositories, is introduced and highlights that coding agent vulnerability is not a static property, but a dynamic outcome shaped by everyday user configurations.
Fu-Kang Zhu, Bin-Bin Zhao, Rui-Xiao Lin et al.· 0 citations
Experiments across multiple datasets, victim models, and large language models show that OASIS consistently outperforms strong standalone baselines and simple manually constructed chains, suggesting that attacker composition is not merely an implementation choice, but a practical optimization target for improving hard-...