Long-horizon LLM agents interact with untrusted content, persistent memory, external state, and sensitive tools. Existing analyses often characterize attacks by the number of execution steps between malicious input and a downstream action. We show that temporal remoteness can overstate security separation in stateful a...
Md Jafrin Hossain, Nur Al Hasan Haldar· 0 citations
KITA is presented, a review-to-authorization architecture that keeps the user's personal secret signing key and every threshold signing-key share outside all LLM processes and establishes execution-bound authorization integrity.
Despite the significant efforts devoted to aligning large language models (LLMs) with human values and ensuring safe deployment, recent work has revealed that LLMs remain vulnerable to adversarial jailbreak attacks that can bypass safety guardrails and elicit harmful responses. Many defense methods are proposed to dete...
Quoc le Viet Vo, Trung Le, D. Ranasinghe et al.· 0 citations
Bait-and-Recover is proposed, a weight-level defense that places a bait adapter where attackers read activations and a paired recovery adapter at the subsequent layer that decouples the observation path from the behavior path.
Tian Gao, Zhi-Hui Xie, Yu-Hao Wu et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
One of the biggest risks faced by Software Defined Networks (SDN) is the Distributed Denial of Service (DDoS) attack in which a compromised controller can make an entire network unusable. To address these challenges, we suggest an entropy-guided machine learning framework, called XAI-SDN, for real-time DDoS detection i...
Adeel Ahmad, Ali Akarma, Ahmad Ali et al.· 0 citations
A guard that sits between the agent and any memory backend and withholds records that are revoked or conflict with their replacement is developed, finding that no system enforces revocation by default.
Yi-Da Shen, Kentaroh Toyoda, Alex Leung· 0 citations
Tool-using language agents can delegate and revoke permissions while acting through external services. We show that two authorization histories can have identical current permissions and identical all-pairs reachability yet require opposite decisions after the same direct-edge revocation. We formalize the information n...
The AI-safety version of the threat-model coverage gap is called the AI-safety version the threat-model coverage gap, and it is found that it persists in a current open-weight model, where harms surface in non-English prompts that English benchmarks miss.
Resource-exhaustion attacks against autoregressive vision-language models (VLMs) typically assume unimodal threat models, treating the image branch as the primary optimization surface while holding user-visible prompts fixed. Even recent loop-centric variants remain confined to this single-channel paradigm, leaving the...
Zhaoxiong Ni, Yatie Xiao, Chi-Man Pun et al.· 0 citations
Experimental results show that dual-embedding watermarking can offer state-of-the-art robustness, particularly against translation, while incurring relatively low computational overhead compared with other semantic schemes.
Jonas Schäfer, Cezary Pilaszewicz, Gerhard Wunder· arXiv.org· 0 citations
Neutral Prompting Attack (NPA), a highly stealthy attack paradigm in which semantically benign instructions, such as encouraging imagination and exhaustiveness, increase package hallucination propensity without containing explicit malicious intent is introduced.
Initial findings in two key areas of offensive security: Privacy-Preserving Machine Learning, and network traffic analysis are presented, by presenting initial findings in two key areas of offensive security.
Giovanni Cherubin· The Importance of Being Lear...· 0 citations