Skip to content
Open access

condfair: An R Package for Ability-Conditioned Fairness and Explanation Diagnostics in Automated Scoring.

Aug 2026 · Applied Psychological Measurement · pp. 01466216261484166 · 0 citations · 3 references
Medicine

Abstract

Fairness in automated scoring is typically evaluated with a single global statistic contrasting a focal and reference group-an approach that can either mask a disparity that changes sign across the ability range, or overstate one by conflating it with genuine ability differences between groups (impact). We introduce condfair, an R package that adapts differential item functioning (DIF) logic to automated scoring: it estimates a conditional disparity function across ability levels, tests it with a wild-bootstrap omnibus procedure, decomposes bias into uniform and non-uniform components, and identifies candidate feature-level sources of a detected disparity via conditional SHAP disparity testing. Using the PERSUADE 2.0 essay corpus, we show a global measure can conceal a large, ability-concentrated gender disparity (marginal gap = 0.003; peak conditional disparity = 0.276, p = .001) while overstating an English Language Learner disparity by conflating it with impact (marginal gap = 0.280; conditional bias = 0.045).

Read PDF

Similar papers

Open access Jul 2026

WHEN FAIR AI BECOMES UNFAIR: A COUNTERFACTUAL AUDIT OF POSITIONAL BIAS IN LARGE LANGUAGE MODELS FOR HIRING DECISIONS

Findings indicate that state-of-the-art LLMs can achieve a high degree of demographic neutrality; fundamental artefacts such as positional bias can nonetheless produce severely discriminatory outcomes; and bias auditing must extend beyond demographic parity to interaction artefacts and ecosystem structure.

A. Camargo, Rafaela Silva Figueiredo Camargo · 0 citations
#artificial intelligence Preprint Aug 2026

Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit

Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded rating scale. We show that this endpoint is not identified on the scale that reports it. Each term of the double difference is censored by its own share, so the observed statistic confounds differential preference with differential attenuation: a severity shift common to both responses manufactures an interaction whenever the two censor it unequally, as unequal distances from the bounds make them, exactly where good stimuli place them. We exhibit the failure inside a pre-registered audit of a frozen pedagogy judge, sealed before the first of its 990 calls. The registered primary endpoint, the effect of a stated learner profile on the judge's scaffolding preference, is null: $+0.085$ points (95\% BCa $[-0.167, +0.353]$, $p = 0.684$). The audit's one nominally significant interaction, $+0.378$ ($p = 0.002$), is not identified as preference: a construction containing zero differential preference reproduces 79 to 85\% of it from the observed severity shift and the scale floor alone. We derive the mechanism in closed form and show that its contribution is measurable from an audit's own ratings.

Shu-Yi Fan, Boyuan Deng, Mengyu Xu et al. · 0 citations
Review Aug 2026

Whistle meets the algorithm: the impact of AI-integrated VAR on satisfaction, fairness, and trust in soccer officiating

The study examines whether exposure to AI-integrated Video Assistant Referee (VAR) technology changes fans' perceptions of the VAR officiating system across the dimensions of procedural fairness, trust, and satisfaction. Moving beyond prior scholarship that largely explores human-assisted review mechanisms, the study considers how fans evaluate algorithmically enhanced officiating systems. A two-phase within-subjects experimental design was employed. Phase 1 examined the stimulus (a video featuring artificial intelligence (AI)-integrated VAR) with 50 participants. Phase 2 tested the hypotheses with 503 soccer fans recruited via Prolific. Wilcoxon signed-rank tests assessed differences between pre-stimulus and post-stimulus measures, and analysis of covariance (ANCOVA) models tested robustness after controlling for fandom and demographic covariates. Exposure to AI-integrated VAR improves perceptions of procedural fairness, trust, and satisfaction with VAR, supporting all three hypotheses. These effects remain robust after accounting for baseline attitudes and relevant covariates. Media consumption also shows a modest yet positive association with trust, suggesting that mediated consumption may help shape evaluations of AI-integrated VAR officiating. This study extends Expectation Confirmation Theory to AI-integrated officiating and shows that fan responses to AI-integrated VAR can be understood not only as reactions to decision outcomes but also as evaluative judgments formed through the dynamics between expectation and confirmation regarding the credibility of the decision-making process. The study contributes to the emerging scholarship on AI-assisted consumer behavior in sport and offers practical implications for leagues, governing bodies, and broadcasters seeking greater public acceptance of officiating innovation.

Ryan Chen, Susmit S. Gulavani · 0 citations

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

FairFund-Bench is introduced, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised, indicating that current LLMs robustly reproduce human deservingness evaluations.

Martin Lukk · 1 citation
Jul 2026

What AI Red-Team Evaluations Can and Cannot Prove

This work defines the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and uses it to locate that boundary exactly.

Bandana Kaur · 0 citations
Jul 2026

Effort Matters in Score-Based Admissions: How Retaking and Aggregation Shape Test Scores

Observed standardized test scores are the result of an endogenous process: students strategically allocate effort across multiple retake attempts to improve their outcomes. Because students differ in their ability to make these investments, the interaction between applicant strategy and institutional scoring rules---such as the widely used Single-Sitting and Superscoring policies---can disparately distort observed scores. We develop a strategic framework where students allocate effort in response to different scoring policies. We show that Superscoring---the practice of combining the best section scores across attempts---introduces systematic score inflation through order-statistic selection over noise draws. This degrades signal accuracy and amplifies wealth-based disparities by disproportionately rewarding applicants who can afford repeated testing. Conversely, Single-Sitting---which keeps the best overall score rather than section-level scores---preserves signal fidelity but excludes high-ability students who lack the resources to prepare for all subjects simultaneously. Neither rule uniformly dominates; instead, they force a structural trade-off between statistical precision and fair outcomes. Finally, to address this, we propose three algorithmic interventions which either modify how scores from multiple attempts are combined, or apply a post-hoc correction to observed scores. Using simulations calibrated to 2025 College Board data, we compare standard scoring rules against these proposed interventions.

Christine Ling, Diptangshu Sen, Juba Ziani · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.