Long-term memory can turn untrusted external content into persistent influence over an LLM agent's future decisions, creating the threat of indirect memory poisoning. A successful attack must survive a multi-stage pipeline comprising memory writing, retrieval, and utilization. Existing attacks largely rely on intra-sta...
Chuan-Chao Zang, Jianing Wang, Wenyu Chen et al.· 0 citations
It is shown that feedback-based agents can retain early biases even when later correction is available, and a black-box framework for exploiting this weakness is proposed that operationalizes the three factors as directional-shift, contextual-plausibility, and counterevidence-resilience signals under either limited tar...
Chuan-Chao Zang, Jia-Ning Wang, Wen-Yu Chen et al.· 0 citations
This paper proposes ToolSiphon, a query-only extraction attack that introduces two complementary signals: a target-discriminative signal, implemented through Tool Contrastive Analysis, to steer queries toward the target tool; and a response-grounded factual signal, implemented through Evidence Chained Feedback, to miti...
Chuan-Chao Zang, Jia-Ning Wang, Wen-Yu Chen et al.· 0 citations
Control evaluations reveal that targeted poisoning risk varies across memory operations and motivate stage-aware evaluation and control of LLM-agent memory, showing that targeted poisoning risk varies across memory operations.
Chuan-Chao Zang, Zi-Jian Cao, Xiang-Tao Meng et al.· 0 citations
The first extraction attack designed for this threat model, SPORE decouples the adversarial command from retrieval anchors by persisting the command in short-term memory and emitting semantically pure anchors in tool responses, demonstrating that memory isolation alone is insufficient and call for reexamining tool-side...