The results indicate that Micro-Collaborative Poisoning is not driven by a single dominant poisoned passage, but by the accumulation of weak adversarial signals across retrieved sources, which achieves downstream influence while leaving a weaker explicit poisoning signature than direct poisoning.
Abstract
Retrieval-Augmented Generation (RAG) improves large language models by grounding outputs in external knowledge sources, but this dependency also creates a surface for poisoning attacks. This paper introduces Micro-Collaborative Poisoning, a distributed attack in which a false target claim is divided across multiple locally plausible documents instead of being concentrated in a single malicious passage. We evaluate the attack across 108 RAG configurations by varying dataset, retriever architecture, retrieval depth, database composition, number of poisoned databases, and generator model. The results indicate that Micro-Collaborative Poisoning is not driven by a single dominant poisoned passage, but by the accumulation of weak adversarial signals across retrieved sources. Increasing top-$k$ and poisoning multiple databases make it more likely that these signals will appear together in the retrieved context, while clean database diversity and stronger retrievers can reduce their influence. The document-level poisoning visibility analysis further shows that this threat is difficult to expose through isolated document inspection, since Micro-Collaborative Poisoning achieves downstream influence while leaving a weaker explicit poisoning signature than direct poisoning.
Att2RAG is presented, a double-condition framework for knowledge poisoning attacks on RAG systems that decomposes a successful poisoning event into a retrieval condition and a generation condition, and casts poisoning as maximizing attack success subject to satisfying both conditions.
Zhize Hao· Poster Volume 0008 The 2026...· 0 citations
It is proved that, under an honest-majority assumption and a representation-level separation condition, RAGSentinel exactly recovers a poison-free majority-sized context.
Yueyang Quan, Anjun Gao, Yu Xia et al.· 0 citations
This work presents RAGSieve, which constructs a reference matched to each detection scope, which requires poison labels, a trusted corpus, or training to be constructed.
ToxicRAG is presented, a one-document-per-target attack that expresses misinformation as a coherent knowledge-update narrative that matches or exceeds the strongest evaluated baseline in every combination of dataset--model combinations.
Hao-Zhe Lu, Jia-Qiang Li, Xin-Yuan Zhu et al.· 0 citations
The Tri-Layer Sieve is presented, a middleware defense that sanitizes retrieved evidence through cross-embedding-space clustering with an independent judge model, structural filtering of trigger-payload artifacts, and LLM consistency verification, and exploits a key weakness of retrieval-stage poisoning.
Muhaimin Bin Munir, Akib Jawad Ononto, Nazia Shehnaz Joynab et al.· 0 citations
ActProbe, an internal-state-based framework for detecting and localizing poisoned segments in multi-source LLM inputs, is proposed and remains effective against defense-aware adaptive attacks and can protect black-box APIs through surrogate-based poisoned-segment removal.
Xue Tan, Chang-Hui Wang, Sanrui Yang et al.· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.