Author

David Thompson

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Review Open access 2020

Explainable Reinforcement Learning for Transparent Automation

Reinforcement Learning (RL) is widely used for solving sequential decision-making problems, enabling agents to learn optimal actions through interaction with dynamic environments. However, many RL models function as “black boxes,” making their decisions difficult to interpret—an issue that is especially critical in safety-sensitive domains like healthcare, finance, and autonomous systems. Explainable Reinforcement Learning (XRL) addresses this challenge by providing human-understandable insights into agent behavior, policy decisions, and reward structures. This work reviews pre-2019 XRL approaches, categorizing them into policy explanation, reward decomposition, model transparency, and post-hoc interpretability methods. It highlights the trade-off between performance and interpretability, particularly in complex, high-dimensional environments. A framework is proposed that combines interpretable policies, surrogate models, attention mechanisms, and visualization techniques to enhance transparency without significantly reducing performance. Evaluation metrics such as fidelity, comprehensibility, and consistency are used to assess explanation quality. The analysis shows that hybrid approaches—combining inherent interpretability with post-hoc explanations—offer the best balance between accuracy and transparency. Overall, XRL is essential for building trust in automated systems, with future research focusing on standardized evaluation methods and human-in-the-loop learning.

Michael Anderson, David Thompson · 0 citations