When the Audit Depends on the Auditor: Prompt Sensitivity in Matched-Pair Audits of Large Language Models
Abstract
Bias audits of large language models (LLMs) typically use a single prompt to compare model responses to identity-matched stimuli. This design rests on the implicit assumption that the estimated identity gap is stable across plausible prompt wordings. We test this assumption in a factorial experiment on three contemporary LLMs (Qwen, Llama, and GPT4o), which evaluated ten matched-text managerial vignettes under four prompts crossed with gender and name-signalled ethnicity manipulations, yielding 48,000 model calls. Prompt choice moved both rating levels and, in some cells, the estimated identity gap. Prompt effects on rating levels were often larger than identity effects, especially when a deliberately critical stress-test prompt was included. Because that critical prompt changes the evaluative construct rather than only its wording, we report results with and without it, treating the three construct-preserving prompts as the primary comparison and the critical prompt as an upperbound stress test. Where the gap was non-trivial, sign reversals were rare; the substantive importance of the prompt-induced instability we observed depends strongly on whether the underlying gap is itself non-trivial. We summarise cross-prompt variation with the Audit Disagreement Index (ADI). Single-prompt audits are not necessarily wrong, but they may be incomplete: reporting across multiple prompts allows readers to distinguish near-null gaps from substantively meaningful and stable ones.