Skip to content
Open access

Imagining Trajectories for Anomaly Detection in Reinforcement Learning from Images

Aug 2026 · Machine-mediated learning · Vol 115 · 0 citations · 31 references

Abstract

Detecting anomalous inputs is a critical prerequisite for the safe deployment of reinforcement learning (RL) agents in real-world environments. Although anomaly detection (AD) has been extensively studied in other domains, its application to reinforcement learning remains challenging due to high-dimensional sensory observations and complex temporal dependencies. Existing approaches in this setting are limited and often rely on access to internal representations of trained agents, creating an undesirable coupling between policy and safety mechanisms. In this work, we propose ITRM, a novel approach to anomaly detection in visual reinforcement learning that is fully agent-agnostic and does not require access to policy internals. Our method is based on the observation that deviations from nominal environment dynamics can be identified through discrepancies in the latent representations of a learned world model. Specifically, we leverage predictive components of a recurrent state-space model to generate deterministic latent embeddings that serve as normative references for anomaly detection. Anomalies are detected by comparing predicted latent features against a reference set of nominal embeddings using a similarity-based criterion. Operating in the world model’s latent space allows the detector to capture semantically meaningful deviations. Extensive empirical evaluations and ablation studies demonstrate that the proposed approach achieves strong detection performance. On the Anomaly-Gym benchmark, our method outperforms existing baselines, achieving an average AUROC of 0.853 and an FPR95 of 0.279.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.