Skip to content
Preprint

Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning

Aug 2026 · 0 citations · 32 references
Computer Science

Abstract

Pareto Conditioned Networks learn multiple multi-objective reinforcement learning behaviours by conditioning a single policy on a desired return command. However, the local mapping from command and state to action remains opaque. We propose command-space counterfactual explanations for PCNs: given a fixed state, original command, and foil action, we search, in a black-box setting, for a minimally changed desired-return command under which the same trained policy would choose the foil. Our contributions are threefold. First, we formulate PCN explanations as return-command interventions, using a return-only PCN variant that avoids the added ambiguity of horizon-conditioning. Second, we adapt adversarial machine learning methods to reinforcement-learning explanations. Third, we introduce a boundary-seeded directional search that improves over purely local optimization in the command-action landscape, resulting in our proposed approach CF-ZOO. The resulting explanations are actionable and intuitively expressed in the user's own preferences:"If your trade-off had shifted slightly towards X, the agent would have chosen Y."

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.