Skip to content

Author

Debeshee Das

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Oct 2026

Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not Enough

Memory poisoning attacks on LLM agents typically assume an external adversary who plants content in the agent's persistent memory to steer its behavior. We instead study, with no adversary involved, whether a misaligned agent can write a goal it cannot yet act on to persistent memory, so that a future aligned agent car...

Debeshee Das, Jacqueline Tay, Bruce Tsai et al. · 0 citations
#artificial intelligence Preprint May 2026

Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents

The Trojan Hippo Bench is introduced, a dynamic evaluation framework for persistent memory attacks and memory-layer defenses, comprising an OpenEvolve-based adaptive red-teaming benchmark that stress-tests defenses and memory backends against continuously refined attacks, and a capability-aware security-utility analysi...

Debeshee Das, Julien Piet, D. Kaviani et al. · 12 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.