Skip to content

Micro-Collaborative Poisoning: A Distributed Attack on RAG Systems

Sep 2026 · 0 citations · 20 references
Computer Science

TL;DR

The results indicate that Micro-Collaborative Poisoning is not driven by a single dominant poisoned passage, but by the accumulation of weak adversarial signals across retrieved sources, which achieves downstream influence while leaving a weaker explicit poisoning signature than direct poisoning.

Abstract

Retrieval-Augmented Generation (RAG) improves large language models by grounding outputs in external knowledge sources, but this dependency also creates a surface for poisoning attacks. This paper introduces Micro-Collaborative Poisoning, a distributed attack in which a false target claim is divided across multiple locally plausible documents instead of being concentrated in a single malicious passage. We evaluate the attack across 108 RAG configurations by varying dataset, retriever architecture, retrieval depth, database composition, number of poisoned databases, and generator model. The results indicate that Micro-Collaborative Poisoning is not driven by a single dominant poisoned passage, but by the accumulation of weak adversarial signals across retrieved sources. Increasing top-$k$ and poisoning multiple databases make it more likely that these signals will appear together in the retrieved context, while clean database diversity and stronger retrievers can reduce their influence. The document-level poisoning visibility analysis further shows that this threat is difficult to expose through isolated document inspection, since Micro-Collaborative Poisoning achieves downstream influence while leaving a weaker explicit poisoning signature than direct poisoning.

View source

Similar papers

Conference 2026

Att2RAG: A Double-Condition Framework for Knowledge Poisoning Attacks on RAG Systems

Att2RAG is presented, a double-condition framework for knowledge poisoning attacks on RAG systems that decomposes a successful poisoning event into a retrieval condition and a generation condition, and casts poisoning as maximizing attack success subject to satisfying both conditions.

Zhize Hao · 0 citations
#natural language process... Preprint Aug 2026

TRIS: A Tri-Layer Retrieval Integrity Sieve Against Knowledge Poisoning

The Tri-Layer Sieve is presented, a middleware defense that sanitizes retrieved evidence through cross-embedding-space clustering with an independent judge model, structural filtering of trigger-payload artifacts, and LLM consistency verification, and exploits a key weakness of retrieval-stage poisoning.

Muhaimin Bin Munir, Akib Jawad Ononto, Nazia Shehnaz Joynab et al. · 0 citations
#small language model Review Sep 2026

Detecting and Localizing Segment-Level Poisoning in Multi-Source LLM-Agent Inputs

ActProbe, an internal-state-based framework for detecting and localizing poisoned segments in multi-source LLM inputs, is proposed and remains effective against defense-aware adaptive attacks and can protect black-box APIs through surrogate-based poisoned-segment removal.

Xue Tan, Chang-Hui Wang, Sanrui Yang et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.