Jul 2026· INTERNATIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT· Vol 10, pp. 1-9· 0 citations
Abstract
In this paper we present a 30, 000-variant India-context bias audit to compare LLaMA3.2-1B and LLaMA3.1-8B across categories of Gender, Religion, Profession and Region that proposes the India Context Sensitivity Index (ICSI) as a category-weighted fairness metric. The larger model shows an improvement of the aggregate bias score that is statistically significant (Mann-Whitney p < 0.001) a biased-response rate that is significantly lower (χ² = 51.92, p < 0.001, 94.04% unbiased responses) at approximately double the inference latency. The breakdown of results by categories also shows that this overall gain is unevenly split: the 8B remains more sensitive on Gender-based prompts even as it improves on Religion, Profession and Region, underscoring the merit of reporting disaggregated, category-imbued fairness over a single number bias score for India deployed LLMs [9], [17], [23]. The future work will allow the scaling of the benchmark to be done to additional model sizes within the LLaMA family (for instance 3B, 70B) to instead fit a bias-versus-scale curve rather than a two-point comparison (14). Then, the extension of the category set to ‘caste-adjacent’ and intersectional categories (for instance gender × region) that are under-represented in this effort (18), (9). The next direction will be replacing the lexicon-based bias scorer with the LLM-as-judge scorer validated against human annotation, to measure scorer-induced bias in the evaluation pipeline itself (28). Finally, deploying the mitigation variants (few-shot, prompt-engineering) evaluated here as live runtime guardrails and measuring their effect on the ICSI in a closed deployment loop (29), (30).
Findings indicate that state-of-the-art LLMs can achieve a high degree of demographic neutrality; fundamental artefacts such as positional bias can nonetheless produce severely discriminatory outcomes; and bias auditing must extend beyond demographic parity to interaction artefacts and ecosystem structure.
A. Camargo, Rafaela Silva Figueiredo Camargo· Revista de Geopolítica· 0 citations
FairFund-Bench is introduced, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised, indicating that current LLMs robustly reproduce human deser...
It is demonstrated that gender bias undermines both reliability and fairness in LLM-based fake news detection, highlighting the need for bias-aware evaluation and mitigation strategies.
Razieh Chalehchaleh, R. Farahbakhsh, N. Crespi· 0 citations
Benevolence bias is identified and measure, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions, and is easy to diagnose and straightforward to fix.
Yuanzi Li, Jun-Hao Wang, Minghui Liu et al.· 0 citations
The Duckworth-Lewis-Stern (DLS) method has been the international standard for revising target scores in rain-interrupted limited-overs cricket since 1999. Despite over two decades of operational use, no large-scale empirical audit of its prediction bias has been published. We conduct such an audit on 8,150 internation...
A PRISMA 2020-guided systematic literature review draws on 82 studies selected from 493 records retrieved from Scopus and Web of Science and reveals a structural disconnect in the fairness-in-NLP and HCAI governance literature.
Asmae El Moutafail, Khalid Belkhoutout· EPJ Web of Conferences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.