Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whether an embedded instru...
The results suggest that private dense retrieval is already practical for moderately sized, privacy-sensitive corpora when minute-scale latency is acceptable, and formulates a two-round cryptographic protocol that provides both guarantees.
Louis Tremblay Thibault, Sofiane Azogagh, Marc-Olivier Killijian et al.· 0 citations
A personal language agent that acts for its owner across private and shared conversations can learn a fact from one audience and later place it in the context it assembles for another. We study authorization before context across the whole memory lifecycle. Each memory item carries the audience present when it was reco...
A key observation is that reference models serve only to reveal how an example behaves under models not trained on it, and that querying the target model on nearby samples yields the same information, and that querying the target model on nearby samples yields the same information.
Francesco Rita, Jie Zhang, Florian Tramèr· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Decentralized Federated Learning (DFL) eliminates the central aggregation server, reducing the single point of observation that traditional defenses against attacks rely on. As a result, peer-to-peer networks become exposed to malicious updates containing backdoors or semantic poisoning, since such updates can remain c...
Pedro Beltr\'an-L\'opez, Enrique Tom\'as Mart\'inez Beltr\'an, Pantaleone Nespoli et al.· 0 citations
Distributed Denial of Service attacks are a growing threat to network infrastructure, and new techniques, including the use of generative AI, make them harder to detect. Traditional detection systems, such as rule based firewalls, often fail to identify these evolving attack patterns. In this study, we propose a new me...
Aadith Sukumar, Isha Singh, Devershika Mohane et al.· 0 citations
Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model's behavior. In this paper, we study the role played by the adversarial data's input modality, and identify a systematic asymmetry: multimodal LLMs are more likely to follow adversarial instructio...
Jie Zhang, Andrei Baroian, J. V. van Rijn et al.· 0 citations
PrivacySkills is introduced, a controlled framework for evaluating how agents choose among acquisition pathways that provide the same task-relevant value: consulting publicly available personal information, accessing confidential sources, or interacting with the user.
Lucas Biechy, Cédric Eichler, H. H. Arcolezi et al.· 0 citations
In every base and instruction-tuned pair the authors test, instruction tuning strengthens the model's preference for reserved markers, and the gap persists on that channel.
Yan-Yan Zhan, Yun-Ze Song, Meng-Kai Hou et al.· 0 citations
Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, attackers can disguise malicious instructions as ordinary visual guidance within otherwise...
Lingqi Jiang, Jialuo Chen, Jianan Ma et al.· 0 citations
This work presents RedHerring, which inserts certifiably safe decoys that divert verification effort from real vulnerabilities, and shows that its effectiveness does not depend on decoy secrecy.
Kai-Kai Zhang, Zi-Han Zhang, Yu-Chong Xie et al.· 0 citations
Semantic caches reduce LLM serving costs by reusing previously generated answers for semantically similar queries. However, retrieval is based solely on embedding similarity between the incoming query and cached queries. This design enables cache poisoning: an attacker can cache a malicious response under a query with...
Zihan Zhang, Shuangjie Yao, Zesen Liu et al.· 0 citations