Skip to content

When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents

Sep 2026 · 1 citation · 35 references
Computer Science

TL;DR

A targeted prompt-level defense is evaluated and finds that it can reduce memory injection in many settings, but provides limited protection once the persistent memory has been poisoned.

Abstract

Harness design has transformed the development of LLM-based agents by integrating memory, tool use, and runtime control. However, this design also introduces security and privacy risks because malicious instructions from external sources may be written into persistent memory and persist across sessions. To study this risk, we propose PMPA, a Persistent Memory Poisoning Attack against harness-based agents. PMPA embeds malicious instructions into benign external sources and induces the victim agent to write them into persistent memory without directly accessing to the agent framework. Once stored, the poisoned memory can be retrieved in later sessions, triggering additional malicious actions and causing privacy leakage. We evaluate PMPA on OpenClaw and Claude Code across different backbone LLMs, input modalities, and trigger scenarios. Across all settings, PMPA achieves average Injection Success Rate (ISR) and Cross-session Attack Success Rate (C-ASR) of 73.7%/ 55.5% on OpenClaw and 66.9%/ 81.7% on Claude Code, while preserving benign task performance on both systems. We further evaluate a targeted prompt-level defense and find that it can reduce memory injection in many settings, but provides limited protection once the persistent memory has been poisoned.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?

This work proposes Continuous Agents for Injection Threats via Lifelong Yielding Nexus (CAITLYN), an agent-agnostic defense middleware that matches the detection performance of state-of-the-art defenses at lower token overhead than LLM-as-a-judge baselines.

Zi Liang, XiaoYu Xu, Yanyun Wang et al. · 1 citation
Preprint Aug 2026

HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses

Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills, tools, and shared artifacts. However, this capability creates delayed safety risks: attacker-influenced content can cross system boundaries and later affect the execution of a benign request. Existing benchma...

X. Zhang, Yusheng Wang, Yuhao Fei et al. · 1 citation
#artificial intelligence Preprint Sep 2026

The Price of Safety: Benign-Case Utility and Token Overhead of Memory-Poisoning Defenses in LLM Agents

Memory-poisoning defenses for LLM agents are typically evaluated by their ability to prevent attacks. However, the traffic they process is rarely adversarial. The cost of implementing a defense is paid with each interaction, while its benefits are only seen in a small percentage of cases. We developed a measurement set...

Pritom Bhowmik · 0 citations
Preprint Aug 2026

SynChain: Inducing Computer-Use Agent Systems to Construct Their Own Attack Chains

This work introduces SynChain, a self-synthesized attack paradigm utilizing persistence-aware directed supervised fine-tuning to induce agents to create poisoned yet benign-looking artifacts, proving that securing CUAs requires provenance-aware reasoning over cross-task execution trajectories.

Fu-Yao Zhang, Jia-Ming Zhang, Che Wang et al. · 0 citations
Preprint Sep 2026

pikit: A Composable Toolkit for Indirect Prompt Injection Research and Evaluation

Indirect prompt injection embeds malicious instructions within external content retrieved by LLM-based agents, altering target behavior without user authorization. We introduce pikit, a research toolkit designed to systematically evaluate these threats across three core dimensions: attacks (13 methods), channels (16 ca...

Zong-Hao Ying, Xiang-Fan Wu, Bo Yang et al. · 0 citations
Open access Sep 2026

Adaptive Honey-Memory: A Semantic Deception Framework for Disrupting Agentic Malware

The rise of agentic AI systems, which are autonomous agents capable of reasoning, planning, and carrying out multi-step tasks, has opened up a previously unknown attack surface: memory-based malware spreading. Traditional intrusion detection systems are not designed to detect attacks that exploit the contextual memory...

Talha Ahsan, Muhammad Asad, Saba Firdous et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.