Adaptive Honey-Memory: A Semantic Deception Framework for Disrupting Agentic Malware
Abstract
The rise of agentic AI systems, which are autonomous agents capable of reasoning, planning, and carrying out multi-step tasks, has opened up a previously unknown attack surface: memory-based malware spreading. Traditional intrusion detection systems are not designed to detect attacks that exploit the contextual memory layers of agents powered by large language models (LLMs). The proactive defense strategy, called Contextual Sabotage, is described, which disrupts the agentic malicious life cycle utilizing Adaptive Honey-Memory (AHM), a dynamically generated semantically realistic yet intentionally deceptive layer of memory that intercepts, confuses, and neutralizes malicious agent behavior. AHM modifies the operational environment that adversarial agents can access to create logical inconsistencies, disrupt execution sequences, and create failure conditions within the attacker's operational environment without divulging any actual memory in the system. The study models the agentic threat model, introduces the honey-memory injection framework and experiments with the proposed approach using simulated adversarial agents under three attack scenarios: context hijacking, propagation of goal misalignment, and permanent memory poisoning. Experimental results show that AHM is able to disrupt the target with a high disruption rate of over 87% on all attack categories, and at the same time, it does not interfere with the legitimate agent operations with a rate of less than 3%. The proposed architecture provides a foundation for semantically aware, deception-enabled protection mechanisms for future autonomous AI settings. Evaluation is based on 1200 trials of adversarial attacks across all three attack classes and 800 trials of benign tasks and reported as mean ± standard deviation over 10 trials per condition.