Skip to content
Open access

Individual and group fairness assessments via counterfactual explanations

Sep 2026 · AI and Ethics · Vol 6 · 0 citations · 50 references

Abstract

This study explores the potential of counterfactual explanations to assess artificial intelligence (AI) fairness, especially in critical decision-making systems. Predictive models may amplify biases inherent in data sets or algorithms, and given the absence of a universally accepted fairness metric, a case-specific approach becomes mandatory. Existing statistical fairness metrics may not capture all aspects that are relevant to a context-aware assessment of non-discrimination. The goal of this work is to define a measure of fairness for AI systems based on explainable artificial intelligence concepts. Specifically, it leverages the analysis of counterfactual explanations of individuals/groups and their comparison with similar individuals/groups. Compared to existing state-of-the-art works, the contributions are (i) extending the definition of individual fairness, not limiting unfairness to decisions based on sensitive attributes but also ensuring similar treatment amongst similar individuals; (ii) revisiting (and generalising) existing notions and introducing new, more refined notions of group fairness based on counterfactuals; (iii) defining quantitative fairness metrics that reflect the evidence gathered through the analysis/comparison of counterfactual explanations.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.