2026· EPJ Web of Conferences· 0 citations· 11 references
TL;DR
A PRISMA 2020-guided systematic literature review draws on 82 studies selected from 493 records retrieved from Scopus and Web of Science and reveals a structural disconnect in the fairness-in-NLP and HCAI governance literature.
Abstract
As large language models emerge as critical infrastructure in labor markets, questions about their governance touch on some of the most contested issues at the intersection of AI development, complex sociotechnical systems and emerging technology regulation. Recruitment tools powered by these models are now widely deployed across industries and they raise urgent concerns about intersectional algorithmic discrimination—patterns of unequal treatment that remain invisible so long as auditors examine only one protected attribute at a time. Trained on historical hiring data, these systems tend to reproduce and in some cases deepen, discriminatory patterns that cut across gender, race and other protected characteristics simultaneously. Two research communities bear on this problem without yet speaking to each other adequately: the fairness-in-NLP literature and the human-centered AI (HCAI) governance literature have each grown substantially, but largely in parallel. To examine where they diverge and why, we conducted a PRISMA 2020-guided systematic literature review drawing on 82 studies selected from 493 records retrieved from Scopus and Web of Science. Grounded theory coding and a concept matrix organize the findings around four thematic axes: bias sources across the hiring pipeline, formal fairness metrics and their mathematical limits, debiasing techniques and governance frameworks. What the concept matrix reveals is, above all, a structural disconnect. Technical studies rarely engage with HCAI design principles governance-oriented work rarely operationalizes the technical limitations that the empirical literature has documented in detail. Five evidence-based design recommendations follow from the analysis; a particularly urgent recommendation concerns the development of non-Western fairness benchmarks. A targeted research agenda addresses intersectional auditing and LLM-specific debiasing as the two highest priority open problems.
FairFund-Bench is introduced, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised, indicating that current LLMs robustly reproduce human deservingness evaluations.
Recruitment processes play a central role in shaping access to employment and social mobility. The increasing use of artificial intelligence in these processes is beginning to change how candidates are evaluated, raising questions about whether regulatory frameworks are sufficient to address emerging challenges. Current regulatory approaches remain focused on the technical aspects of AI, while giving less weight to the social conditions that shape access to employment. The article employs qualitative analysis to examine how AI-driven recruitment processes, though often presented as neutral, are influenced by candidates' social and economic backgrounds. It examines how existing patterns of human decision-making can be reproduced through AI systems, particularly as these models rely on historical data and extend existing selection practices. The analysis draws on qualitative data on retraining trajectories and career transitions in the Romanian IT sector. It further shows that access to employment is strongly shaped by differences in access to resources, enabling certain candidates to bypass selection processes more effectively than others. AI systems tend to reinforce these patterns by operating within opaque decision-making structures, thereby contributing to the persistence of unequal outcomes. The findings indicate that screening based on easily measurable criteria is central to hiring practices, whether carried out by human recruiters or AI systems, often limiting candidates' ability to demonstrate their full potential. By bringing together social and technical perspectives, the article highlights how these disadvantages remain embedded in recruitment processes, even when criteria appear neutral. The paper argues for greater regulatory involvement in guiding AI toward more robust and fair evaluation practices and suggests a shift in regulatory approaches toward addressing broader inequalities, enhancing transparency, and promoting more meaningful forms of evaluation in recruitment.
Unknown authors· New Trends in Sustainable Bu...· 0 citations
Findings indicate that state-of-the-art LLMs can achieve a high degree of demographic neutrality; fundamental artefacts such as positional bias can nonetheless produce severely discriminatory outcomes; and bias auditing must extend beyond demographic parity to interaction artefacts and ecosystem structure.
A. Camargo, Rafaela Silva Figueiredo Camargo· Revista de Geopolítica· 0 citations
Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argumentation, and legal reasoning that are difficult to capture with a single metric. In this work, we introduce Poli-Bias, a counterfactual framework for measuring whether LLMs treat legally equivalent conflict scenarios differently depending on the countries involved. Poli-Bias compares responses to paired prompts in which country identities are systematically swapped across diverse geopolitical relationships, legal violations, and reasoning tasks. Rather than reducing bias to a single judgment, our framework decomposes response disparities into five interpretable dimensions, revealing how and where unequal treatment manifests. Across 13 contemporary LLMs spanning diverse model families and sizes, we find that country identities and user affiliations can systematically affect how equivalent actions are described, evaluated, and defended under international law. Our results thus establish Poli-Bias as a fine-grained framework for auditing political even-handedness and sycophancy in LLMs.
Massi-Nissa Abboud, Aladin Djuhera, Elena Cabrio et al.· 0 citations
The rise of large language models (LLMs) has sparked worries about inherent social biases and issues related to fairness. Earlier studies have investigated bias identification in word embeddings, interventions aimed at fairness in algorithms, and frameworks for auditing at the system level. Nonetheless, these methods remain disorganized, with variations in datasets, evaluation methods, and implementation processes. In this paper, we provide a thorough literature review to encapsulate prior research on bias identification and fairness auditing, categorizing the findings according to various stages of study. Additionally, we analyze the limitations in coverage and consistency of widely used benchmark datasets. To tackle these issues, we propose a unified pipeline for dataset integration and a modular framework for bias auditing. Recognized significant research gaps include the absence of intersectional bias modeling, a shortage of standardized evaluation metrics, and challenges in scalability for real-time auditing systems.
Nani Kartik Kaveti, T. Pattanshetti· Discover Artificial Intellig...· 0 citations
AI‐hiring technology is becoming increasingly widely adopted, raising important questions about what effect such technology may have on organizational inequality regimes. Although there is a large and growing body of literature surrounding the effects of AI hiring, its impacts, and bias, markedly less is known about who develops, sells, and markets such tools, and what role this may play in the altering and further entrenching of inequality regimes. Drawing on the voices of 21 developers and hiring platform vendors, this study explores current levels of apathy surrounding gender equality in AI‐hiring tech. From the findings, it is argued that DEI issues and subsequent debiasing are seen as iterative, retrofit solutions, problems that can be tended to retrospectively, rather than inclusion and equity being built‐in from the outset. In turn, it is posited that algorithmic bias and AI‐hiring tools serve as a contemporary facet of inequality regimes. This article offers a threefold contribution to Feminist AI scholarship: a theoretical extension to Acker's inequality regimes framework in the form of an annex, which includes the role of artificial intelligence in the reproduction of inequalities, a conceptual model which outlines new accelerants of inequality regimes, and contemporary empirical insights from a hard to access group.
Emily Yarrow· Gender, Work & Organizat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.