Indirect prompt injection attacks - malicious instructions embedded in content processed by large language models - remain a major obstacle to safely deploying tool-using agents. CaMeL [Debenedetti et al., 2025] mitigates this threat for an individual agent by separating trusted control flow from untrusted data and enf...
James Peters-Gill, Avi Semler, Henning Bartsch et al.· 0 citations
We study state-adversarial Markov decision processes (SA-MDPs) as games of observation-space attacks: at each step, an agent selects an action from a received observation while an adversary$\unicode{x2014}$who knows the true state the agent is in$\unicode{x2014}$chooses a perturbed observation within a state-dependent...
Multi-agent LLM systems are increasingly deployed for complex, long-horizon tasks or emerge as a natural consequence of agents interacting in the wild. Yet they give rise to significant safety and security risks: the flexible protocols that enable task generalization also expose novel threats, from cascading prompt inj...
Ben Hagag, William L. Anderson, S. Chakraborty et al.· 0 citations
A long-held intuition in interpretability research is that representational entanglement, the sharing of structure between knowledge domains in a neural network, makes unlearning harder. While the intuition is widespread, it has never been directly tested in a controlled experiment. We present a way to do so: by repurp...
Evžen Wybitul, Tim G. J. Rudner, C. D. de Witt· 0 citations
It is argued that machine learning-based climate models should be designed to pass generalization tests to prevent overfitting on present-day regional climate.
Maren Höver, Milan Klöwer, Christian Schroeder de Witt et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.