Skip to content

Explaining Reinforcement Learning Agents via Inductive Logic Programming

Jul 2026 · arXiv.org · Vol abs/2607.13655 · 0 citations · 59 references
Computer Science

TL;DR

This work employs Inductive Logic Programming (ILP) to extract symbolic representations of RL policies and define a novel set of explainability metrics, including activation rate, feature coverage, syntactic distance and semantic distance, which provide crucial insights for the transfer and generalization of action-specific policies.

Abstract

Explainable Reinforcement Learning (XRL) seeks to make Reinforcement Learning (RL) policies more transparent and interpretable, a key requirement in safety-critical and human-centric scenarios. However, it is mostly based on user studies, thus targeting the needs of a specific audience and lacking shared evaluation metrics. On the other hand, logic-based approaches within eXplainable Artificial Intelligence (XAI) provide compact, human-readable abstractions of decision-making. However, the systematic quantification of the explainability degree of logical representations remains an open problem. This work aims to advance the state of the art in XRL by introducing objective and planning-oriented metrics for policy explainability in RL settings. At the same time, it contributes to the field of logic for XAI by providing a principled way to quantify the explainability of logical rules, moving beyond common-sense assessments and simple propositional fragments. We employ Inductive Logic Programming (ILP) to extract symbolic representations of RL policies and define a novel set of explainability metrics, including activation rate, feature coverage, syntactic distance and semantic distance. These metrics quantify alignment between symbolic rules and agent behavior, the role of features in decision-making, and the evolution of policies during training and across agents in single and multi-agent RL. Experiments across different RL domains show that the proposed metrics highlight action-specific learning dynamics beyond global return, provide fine-grained insights into domain features beyond classical approaches for global feature importance estimation, and uncover coordination, specialization, and adaptation patterns in MARL. Moreover, they provide crucial insights for the transfer and generalization of action-specific policies.

View source

Similar papers

Review Open access 2020

Explainable Reinforcement Learning for Transparent Automation

This work reviews pre-2019 XRL approaches, categorizing them into policy explanation, reward decomposition, model transparency, and post-hoc interpretability methods, and proposes a framework that combines interpretable policies, surrogate models, attention mechanisms, and visualization techniques to enhance transparency without significantly reducing performance.

Michael Anderson, David Thompson · 0 citations
Review Jul 2026

A survey on explainable reinforcement learning: state of the art, challenges and opportunities

From this literature review, it becomes evident, in face of a series of preliminary and promising studies, the field of XRL still lacks a proper accounting towards full explainability and it emerges that more effort should be devolved into developing paradigms with a human-in-the-loop factor and standardized metrics should be adopted to allow a fair comparison of different explanation methods.

Daniele Melloni, Andrea Zingoni · 0 citations
Preprint Jul 2026

Explaining Reinforcement Learning Decisions in Self-adaptive Systems

Reinforcement Learning (RL) has been extensively used in autonomous and self-* systems, but RL policies, especially deep RL ones relying on neural networks, lack transparency and are difficult to understand. This can lead to diminished user trust, and makes for a more challenging verification of systems. To address this challenge, this paper introduces Explanations using Alternative Realities for Reinforcement Learning (EARL), a Python library to produce counterfactual explanations in RL settings. This library allows the user to produce explanations by exploring What-if scenarios to clarify agent behavior by comparing possible outcomes. Counterfactual explanations have been shown to be intuitive and user-friendly in psychology research, but have only recently been explored in RL, with existing implementations usually limited to toy examples and benchmarks. EARL supports counterfactual explanation generation in realistic RL-based self-adaptive systems. To demonstrate its applicability, we demonstrate its use in a simulation of CitiBikes, a self-adaptive bike-sharing system, and we provide evaluations showing how it performs in real applications.

Jasmina Gajcin, Juan C. Rosero, Ivana Dusparic · 0 citations

Dissecting Reinforcement Learning: Mechanisms Behind Compositional Reasoning in LLMs

This thesis proposes a unified two-axis framework that organizes SFT and RL methods along a data axis (off-policy to on-policy) and a loss function axis (positive-only to positive-plus-negative to GRPO) and enables controlled ablations of individual components.

G. Kim, Chair Chenyan Xiong, Aditi Raghunathan · 0 citations

Declarative Specifications for Efficient and Safe Reinforcement Learning

This dissertation presents a work in safe RL, where agents must also respect safety constraints using pure-past linear-time temporal logic (PPLTL), and presents how to enforce safety constraints using pure-past linear-time temporal logic (PPLTL).

Giovanni Varricchione · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.