Skip to content
Open access

Reassessing demographic bias in face attribute classification: a statistically grounded multi-model evaluation on FairFace and UTKFace

Aug 2026 · Frontiers in Artificial Intelligence · Vol 9 · 0 citations · 25 references
Medicine

TL;DR

A statistically grounded evaluation of demographic bias in face attribute classification across three representative architectures, ResNet50, MobileNetV3, and a vision transformer, using the FairFace and UTKFace datasets shows that race-based disparities are substantially larger than gender-based disparities across both datasets.

Abstract

Face analysis systems are widely used in security, authentication, and public-sector applications; however, demographic bias and the statistical reliability of reported performance remain key concerns. Many studies rely on aggregate accuracy without quantifying subgroup disparities or uncertainty, potentially overstating model fairness. This study presents a statistically grounded evaluation of demographic bias in face attribute classification across three representative architectures, ResNet50, MobileNetV3, and a vision transformer (DeiT), using the FairFace and UTKFace datasets. Subgroup analysis is conducted across race and gender, incorporating disparity indices, bootstrap confidence intervals, and inferential statistical testing with effect size analysis. The evaluation uses an embedding-based nearest-neighbor approach to examine representation-level behavior consistently across models. Results show that race-based disparities are substantially larger than gender-based disparities across both datasets. On FairFace, race disparity gaps range from 0.1124 to 0.1266, while on UTKFace they increase significantly to 0.4726–0.4944, with large effect sizes (Cohen's d>1). In contrast, gender disparities remain smaller, with gaps between 0.0280 and 0.0582 on FairFace and 0.0194–0.0326 on UTKFace, and correspondingly small effect sizes (d < 0.13). Despite modest differences in overall accuracy across models, subgroup disparities remain statistically significant across all architectures. These findings emphasize the importance of subgroup-level evaluation, uncertainty quantification, and statistical validation for reliable fairness assessment in face analysis systems.

Read PDF

Similar papers

Preprint Aug 2026

Bias Mitigation in Face Recognition via Demographic-based Supervised Contrastive Learning

Face recognition systems have been shown to be biased toward certain demographic groups by exhibiting different error rates across gender, age, or ethnicity. Though the imbalance of the training data with respect to these demographics is one cause of this bias, training on artificially balanced groups does not completely mitigate the problem. For deployment, face recognition typically works at operating points allowing very low false match rates and, hence, on the tail of the non-match score distribution. While class balancing can improve the means of these distributions, the aim of our approach is to improve fairness by addressing the behavior in the tail. Particularly, we propose the Demographic-based Supervised Contrastive loss (DeSCon) for face recognition, which relies on a well-designed composition of training batches and demographic-aware pair selection. Our experimental evaluation on both demographically-labeled datasets and standard verification benchmarks shows that DeSCon can improve fairness beyond balancing training datasets while maintaining competitive verification performance. Source code is available upon request.

Yu Linghu, Salman Mohammad, Xin-Yi Zhang et al. · 0 citations
Preprint Aug 2026

CIFA: Contextual-Intersectional Fairness Auditing for Hidden Subgroup Discovery in Face Analysis

Fairness evaluation in computer vision commonly relies on aggregate accuracy and demographic subgroup analysis. However, visual models are also sensitive to contextual factors such as illumination, blur, image quality, facial accessories, and appearance attributes. These factors may interact with demographic characteristics, producing hidden subgroups in which performance degrades substantially despite strong aggregate accuracy and apparently acceptable demographic fairness. To address this, we propose the Contextual-Intersectional Fairness Auditing Framework (CIFA), a structured framework for identifying subgroup vulnerabilities arising from interactions between demographic and contextual attributes. CIFA performs demographic, contextual, and contextual-intersectional auditing, followed by worst-group discovery to identify and rank the most vulnerable attribute combinations. We evaluate CIFA on gender classification using ResNet-50 \cite{he2016deep} and ViT-B/16 \cite{dosovitskiy2020image} across FairFace \cite{Karkkainen2021}, CelebA \cite{Liu2015}, and UTKFace \cite{Zhang2017}. Our results show that aggregate accuracy and demographic-only evaluation can mask substantial contextual-intersectional disparities. We further assess several established mitigation strategies through an audit--mitigate--reaudit protocol and find that, although some worst-group disparities are reduced, no single strategy consistently eliminates them across datasets and architectures. These findings establish contextual-intersectional auditing as an important component of fairness evaluation and provide a reproducible framework for discovering, prioritizing, and reassessing hidden subgroup risks in face analysis systems.

Nazia Aslam, Khalid Alsayed, T. Moeslund et al. · 0 citations
Open access Aug 2026

EquiAI: A Vision Foundation Model for Fair and Robust Cross-Demographic Remote Identity Verification

EquiAI is proposed, a robust fairness-aware remote identity verification framework that integrates a pretrained Vision Foundation Model (DINOv2), fairness-aware representation learning, adaptive feature alignment, presentation attack detection, and explainable artificial intelligence into a unified architecture.

Suman Kumar Sanjeev Prasanna · 0 citations
Conference Jul 2026

Fairness-Aware Multi-Task Learning for Age-Gender Prediction with Demographic Parity Constraints

Accurate forecasting of gender and age from facial image recognition is a critical mission benchmarked in person-machine interface, security, and personalized services. Usually, older-style single-task models may fail to draw in any manner on the relations between the attributes of age and gender, falsely predicting outcomes across demographic groups. This work then envisages the two new fairness-aware hybrid models for multi-task learning, DP-FairHybrid-MTL and EquiAgeGen-HybridNet. These models will bring together convolutional neural networks, attention mechanisms, and integrate the whole-scale facial root and local feature-extracting characteristics whilst ensuring demographic parity. As observed by way of benchmarking on facial data-sets, DP-FairHybrid-MTL logs a 94.8% gender-accuracy rate on the gender F1-score, 0.946, and on the mean absolute error, MAE=3.21 years of age, beneficially distinguishing the MT and ST baselines. The further improvements, performance-wise, will go to EquiAgeGen-HybridNet, securing 95.6% for gender-accurate, 3.05 years on the MAE, and an F1-score of 0.952, while consistently maintaining parallel predictions for demographic groups. These results stress that fairness-aware hybrid multitask learning improvements with predictability are influential, due to the fact that its mixed methods help tackle the bringing together of problems related to bias and privilege, grounding a reliable and robust frame for ethically founded and highly validated social analysis on faces.

Bharti Saxena, R. Chaure, Ritu Shrivastava · 0 citations
Preprint Aug 2026

UFPR-PEs: A Brazilian Face Recognition Benchmark with Self-Declared Race/Color Labels

While face recognition systems are widely deployed, ensuring their demographic reliability and robustness under uncontrolled visual conditions remains a critical challenge. To bridge this gap, we present UFPR-PEs, a benchmark for face recognition bias evaluation using public videos of elected Brazilian politicians annotated with official self-declared race/color categories. The dataset adopts the Brazilian census taxonomy, including the parda category, which has no direct equivalent in the U.S.- or Europe-centric schemas commonly used in prior benchmarks. Our benchmark is built from compressed public video and preserves difficult samples so that performance can be analyzed under realistic conditions. We describe the construction pipeline, report dataset statistics, and evaluate face recognition performance across verification and (closed- and open-set) identification settings, including subgroup analysis by race/color and difficulty level. The results show that recognition performance varies substantially with image quality, and that subgroup gaps must be interpreted jointly with visual difficulty rather than in isolation. Overall, UFPR-PEs provides a reproducible and demographically grounded setting for studying face recognition bias under challenging public video conditions.

Alexandre Diano, Bernardo Biesseck, Gabriel Polo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.