Skip to content
Open access

Measuring Bias In Large Language Models: A Comparative Evaluation of LLaMA3.2-1B and LLaMA3.1-8B Across Indian Socio-Culture Dimension

Jul 2026 · INTERNATIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT · Vol 10, pp. 1-9 · 0 citations

Abstract

In this paper we present a 30, 000-variant India-context bias audit to compare LLaMA3.2-1B and LLaMA3.1-8B across categories of Gender, Religion, Profession and Region that proposes the India Context Sensitivity Index (ICSI) as a category-weighted fairness metric. The larger model shows an improvement of the aggregate bias score that is statistically significant (Mann-Whitney p < 0.001) a biased-response rate that is significantly lower (χ² = 51.92, p < 0.001, 94.04% unbiased responses) at approximately double the inference latency. The breakdown of results by categories also shows that this overall gain is unevenly split: the 8B remains more sensitive on Gender-based prompts even as it improves on Religion, Profession and Region, underscoring the merit of reporting disaggregated, category-imbued fairness over a single number bias score for India deployed LLMs [9], [17], [23]. The future work will allow the scaling of the benchmark to be done to additional model sizes within the LLaMA family (for instance 3B, 70B) to instead fit a bias-versus-scale curve rather than a two-point comparison (14). Then, the extension of the category set to ‘caste-adjacent’ and intersectional categories (for instance gender × region) that are under-represented in this effort (18), (9). The next direction will be replacing the lexicon-based bias scorer with the LLM-as-judge scorer validated against human annotation, to measure scorer-induced bias in the evaluation pipeline itself (28). Finally, deploying the mitigation variants (few-shot, prompt-engineering) evaluated here as live runtime guardrails and measuring their effect on the ICSI in a closed deployment loop (29), (30).

Read PDF

Similar papers

Open access Jul 2026

WHEN FAIR AI BECOMES UNFAIR: A COUNTERFACTUAL AUDIT OF POSITIONAL BIAS IN LARGE LANGUAGE MODELS FOR HIRING DECISIONS

Findings indicate that state-of-the-art LLMs can achieve a high degree of demographic neutrality; fundamental artefacts such as positional bias can nonetheless produce severely discriminatory outcomes; and bias auditing must extend beyond demographic parity to interaction artefacts and ecosystem structure.

A. Camargo, Rafaela Silva Figueiredo Camargo · 0 citations

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

FairFund-Bench is introduced, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised, indicating that current LLMs robustly reproduce human deser...

Martin Lukk · 1 citation
Review Jul 2026

Analyzing and Correcting Benevolence Bias in Large Language Models

Benevolence bias is identified and measure, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions, and is easy to diagnose and straightforward to fix.

Yuanzi Li, Jun-Hao Wang, Minghui Liu et al. · 0 citations
#machine learning Preprint Sep 2026

A Fairness Audit of the Duckworth-Lewis-Stern Method: Format-Specific and Gender-Differential Bias, with an Interpretable Calibration Layer for Cricket Target Revision

The Duckworth-Lewis-Stern (DLS) method has been the international standard for revising target scores in rain-interrupted limited-overs cricket since 1999. Despite over two decades of operational use, no large-scale empirical audit of its prediction bias has been published. We conduct such an audit on 8,150 internation...

Soumyadeep Roy · 0 citations
Conference Open access 2026

Bias and Fairness in LLM-Based Recruitment: A Systematic Review

A PRISMA 2020-guided systematic literature review draws on 82 studies selected from 493 records retrieved from Scopus and Web of Science and reveals a structural disconnect in the fairness-in-NLP and HCAI governance literature.

Asmae El Moutafail, Khalid Belkhoutout · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.