2026· International Conference on Data Technologies and Applications· pp. 1155-1162· 0 citations· 32 references
Computer Science
TL;DR
A robust, reliable, and reproducible methodology for evaluating group fairness in classification algorithms, and provides a methodological approach for assessing fairness within the Sufficiency criterion by operationalizing Calibration.
Abstract
: This paper proposes a robust, reliable, and reproducible methodology for evaluating group fairness in classification algorithms. Building upon established theoretical definitions of non-discrimination criteria - Independence, Separation, and Sufficiency-we present a comprehensive approach to quantifying fairness. Notably, our methodology extends beyond binary group comparisons to accommodate scenarios with multiple sensitive groups. Furthermore, we provide a methodological approach for assessing fairness within the Sufficiency criterion by operationalizing Calibration, elucidating critical issues and conceptual subtleties that appear to have been overlooked in existing literature. We assess the reliability and robustness of our model through applications to one real dataset.
Abstract Background Fairness evaluation is essential for trustworthy clinical risk prediction. However, existing fairness-oriented discrimination metrics either ignore cross-group comparisons or rely on exhaustive pairwise evaluations, making them difficult to interpret and impractical for model selection. Objective This study aimed to develop and evaluate novel fairness-oriented discrimination metrics for clinical risk prediction that address limitations of within-group and pairwise cross-group approaches. Methods We examined theoretical properties of existing U-statistic–based metrics, including concordance index (CI) and area under the receiver operating characteristic curve (AUC), when applied to subgroups. We highlighted the distinction between within-group discrimination (ranking within a subgroup) and group-level discrimination (ranking relative to the broader population). Building on this framework, we proposed group-level extensions of the CI and AUC that summarize subgroup-specific performance in a single interpretable measure. We then applied these metrics to the PREVENT (Predicting Risk of Cardiovascular Disease Events) equation, a recently developed model for atherosclerotic cardiovascular disease. Results The traditional subgroup-specific CI and AUC captured within-group but not group-level discrimination, obscuring inequities in clinical decision-making. Existing cross-group approaches (eg, the xCI and xAUC metrics) addressed this limitation but became computationally and interpretively burdensome with multiple subgroups due to pairwise comparisons. Our proposed metrics provided a streamlined alternative, yielding 1 summary statistic per subgroup while retaining sensitivity to cross-group ranking disparities. Applied to PREVENT, these metrics revealed differences in subgroup performance not apparent from within-group evaluations. Conclusions By distinguishing between within-group and group-level discrimination, our framework clarifies a common source of misinterpretation in fairness evaluation. The proposed group-level extensions of the CI and AUC provide practical, interpretable tools for evaluating fairness in clinical prediction models, enabling more transparent and equitable risk assessment.
Haoyuan Wang, Chuan Hong, Michael J. Pencina et al.· JMIR AI· 0 citations
This work introduces a method to lower-bound the discrepancy of a classifier: a quantity that jointly captures inaccuracy and unfairness, and develops a computationally efficient procedure for calculating the tightest possible lower bound on the classifier’s discrepancy.
Algorithms shape high-stakes decisions across society. While promising efficiency, algorithms also raise fairness concerns. This article proposes that algorithmic fairness is best understood as a human-technology interaction problem rather than a purely technical challenge. Algorithms can reproduce human biases, amplify them through feedback loops, or create new forms of unfairness through objectives, proxies, and seemingly neutral variables. Yet they can make decision processes more explicit, disparities more visible, and actively mitigate discrimination. Fairness depends not only on statistical properties but also on how algorithms are designed, used and experienced by those affected by their decisions. This article therefore offers an interdisciplinary perspective that integrates insights from social justice, psychology, computer science, judgment and decision-making, and management.
Unknown authors· Current Opinion in Psycholog...· 0 citations
Neighborhood-based fairness audits evaluate individual fairness by comparing predictions among similar individuals in feature space. Despite their widespread use, little is known about the robustness of the auditing procedure itself. Because these audits rely on nearest neighbor relationships, small perturbations in feature space can alter local neighborhoods and produce different fairness assessments even when model predictions remain unchanged. We develop a geometric framework for analyzing the robustness of neighborhood-based fairness audits under bounded perturbations. Our analysis establishes sufficient conditions for neighborhood invariance, quantifies how neighborhood replacement propagates to audit instability, and introduces audit volatility, a measure of the expected sensitivity of fairness audits under repeated perturbations. Experiments on benchmark datasets support the theoretical analysis and show that the proposed framework explains the observed stability of neighborhood-based fairness audits.
As regulatory requirements increasingly shape automated lending decisions, fairness remains a critical challenge in high-stakes domains, particularly credit scoring. Although artificial intelligence models can achieve strong predictive performance, they may also reproduce biased outcomes that reduce financial inclusion or transfer harm to overlooked protected groups. Existing fairness interventions commonly operate at a single stage of the decision-making pipeline, despite bias often propagating across representational and decision layers. This study proposes iCert-Fair, a two-layer framework for technical fairness assessment and harm recovery in credit scoring. The first layer adopts a fairness-through-explainability paradigm, using SHAP-based explanations to identify direct and proxy dependence on protected attributes and guide structural dataset repair, while the second layer applies targeted threshold-policy adjustments to recover residual harm while preserving decision utility. Experiments on the German and Taiwanese credit datasets show that fairness gains are model- and dataset-specific and may be collective, concentrated, transferred, or recovered unevenly across protected attributes. The direct comparison with representative pre-processing, in-processing, and post-processing methods revealed that baseline methods targeting one protected attribute at a time frequently transferred residual harm to other monitored attributes. In contrast, the fairness-focused recommendations generated by iCert-Fair achieved larger collective fairness improvements across all considered protected attributes while avoiding residual harm. These gains were obtained while preserving predictive utility on the German dataset and with utility degradation remaining below 5% across the evaluated performance metrics on the Taiwanese dataset, alongside consistently lower false-negative risk. The empirical findings support the use of complementary structural and policy-level interventions and demonstrate the importance of jointly evaluating aggregate disparity, worst-case attribute-level harm, cross-attribute transfer, and predictive utility.
The rise of large language models (LLMs) has sparked worries about inherent social biases and issues related to fairness. Earlier studies have investigated bias identification in word embeddings, interventions aimed at fairness in algorithms, and frameworks for auditing at the system level. Nonetheless, these methods remain disorganized, with variations in datasets, evaluation methods, and implementation processes. In this paper, we provide a thorough literature review to encapsulate prior research on bias identification and fairness auditing, categorizing the findings according to various stages of study. Additionally, we analyze the limitations in coverage and consistency of widely used benchmark datasets. To tackle these issues, we propose a unified pipeline for dataset integration and a modular framework for bias auditing. Recognized significant research gaps include the absence of intersectional bias modeling, a shortage of standardized evaluation metrics, and challenges in scalability for real-time auditing systems.
Nani Kartik Kaveti, T. Pattanshetti· Discover Artificial Intellig...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.