Skip to content

One Example Is Enough to Pass Fairness Benchmarks: Rethinking Fairness Evaluation for Aligned LLMs

Sep 2026 · 0 citations · 42 references
Computer Science

TL;DR

It is argued that BBQ-style multiple-choice abstention benchmarks measure a single structural cue, and a model that solves them does not thereby become fair, and is called for evaluation suites that cover a broader spectrum of fairness alignment.

Abstract

Warning: This submission studies stereotypes and biases, and contains toxic and offensive examples, used for illustration purposes only. Fairness benchmarks such as BBQ have become the de facto standard for fairness evaluation across major model families. We argue that these benchmarks are too easy to support their role: training Qwen 2.5 7B Base with Group Relative Policy Optimization (GRPO) on a single BBQ example, or placing that example in context as a one-shot demonstration for in-context learning (ICL), lifts mean BBQ accuracy from 79.9% to 92.9% and 99.0%, respectively, closing 80% of the gap to its large-scale RLHF counterpart (96.1%) with GRPO, and surpassing it with ICL. These effects generalize across model families. A cross-conditioning analysis shows the improvement is carried by the reasoning traces generated by the model, and one example suffices to elicit a category-agnostic ``missing evidence''reasoning pattern. We argue that BBQ-style multiple-choice abstention benchmarks measure a single structural cue, and a model that solves them does not thereby become fair. We call for evaluation suites that cover a broader spectrum of fairness alignment.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Fair Like Us? Auditing LLM Alignment in Resource Allocation

Fair allocation of scarce, indivisible resources is an important challenge in many societal problems. While there are several formal theories of fairness, no single definition can always be satisfied. As large language models (LLMs) are increasingly used to support decisions and act as agents, they raise new concerns a...

Qi-Shen Han, Hadi Hosseini, Joshua Kavner et al. · 0 citations
#machine learning Preprint Sep 2026

When Post-Processing Fairness Constraints Help and When They Harm: Evidence from Eight Cross-Domain Evaluations

Fairness audits in production ML typically occur once, at deployment, on a single domain. Both fail in practice: fairness can shift after retraining or a changing user base, and interventions validated on one dataset are rarely tested across the heterogeneous domains an organization deploys. We present FAPE (Fairness A...

Nithin Raghava Ramachandra Narla · 0 citations
Preprint Aug 2026

Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales

It is demonstrated in controlled experiments that norm-breaking fine-tuning yields norm-divergent actions justified by self-interested rationales, suggesting a systematic shift in patterns of justification.

Long Hoang Nguyen, Brice Valentin Kok-Shun, Guangyu Du et al. · 0 citations
Open access Sep 2026

Individual and group fairness assessments via counterfactual explanations

This study explores the potential of counterfactual explanations to assess artificial intelligence (AI) fairness, especially in critical decision-making systems. Predictive models may amplify biases inherent in data sets or algorithms, and given the absence of a universally accepted fairness metric, a case-specific app...

Federico Sabbatini, Roberta Calegari · 0 citations
#natural language process... Preprint Sep 2026

SCM-based Fairness and Faithful Explainability for Legal Document Classification

Transformer models such as LegalBERT are increasingly used in legal decision support, raising concerns about both fairness and the transparency of model explanations. These properties are usually evaluated separately, leaving open whether a debiasing intervention that changes fairness also changes how faithfully explan...

Yasmina El Kacemi, S. S. M. Ziabari, A. M. M. Alsahag · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.