The Impact of Model Selection Metrics during Hyperparameter Tuning on Algorithmic Fairness: An Empirical Study
The growing use of machine learning in high-stakes domains raises concerns about fairness. The role of optimization metrics in shaping these outcomes remains underexplored. Using a controlled setup, this study investigates how seven performance metrics used for hyperparameter tuning and model selection affect fairness outcomes across five benchmark datasets. Results show that metrics are not neutral: recall-based optimization yields higher disparities, while precision and specificity lead to more balanced outcomes, with PR-AUC showing intermediate behavior. Overall, metric choice influences fairness, but outcomes are largely driven by dataset characteristics, with optimization redistributing errors rather than eliminating bias.