Skip to content

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

Jul 2026 · arXiv.org · Vol abs/2607.26981 · 0 citations · 34 references
Computer Science

TL;DR

This work introduces OptimismBench, which detects directional bias with inverted pairs: each scenario elicits both P(success) and P(failure), and asymmetry between the two framings yields a signed bias score without ground truth.

Abstract

Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a systematic directional tilt has been hard to detect: calibration metrics aggregate unsigned errors, and naturalistic uncertainty offers no ground-truth probability. When an LLM rates a startup's success at 70% but its failure at 15%, the missing 15 points expose a distortion no aggregate score flags. We introduce OptimismBench, which detects directional bias with inverted pairs: each scenario elicits both P(success) and P(failure), and asymmetry between the two framings yields a signed bias score without ground truth. Across 16 models from 8 providers, fourteen are optimistic; pessimism appears only in Anthropic's frontier tier. Eleven matched base-versus-chat pairs across four families show post-training sets the sign of the bias, with opposite shifts in different families. The pattern survives prompt, temperature, perspective, and self-debiasing ablations. A seventeen-model six-language comparison further shows model identity dominates language, with inter-model variance at 4.7x inter-language variance. We release 3,870 items across 10 languages for per-model directional-bias auditing. When alignment makes a model more helpful, it also tilts its probabilities; downstream pipelines inherit the tilt by default.

View source

Similar papers

Review Jul 2026

Analyzing and Correcting Benevolence Bias in Large Language Models

Benevolence bias is identified and measure, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions, and is easy to diagnose and straightforward to fix.

Yuanzi Li, Jun-Hao Wang, Minghui Liu et al. · 0 citations
Preprint Aug 2026

When Text and Numbers Disagree: Evidence Arbitration in Large Language Models

Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence. We study how LLMs arbitrate between such sources when they support opposing decisions. To do so, we introduce a controlled synthetic benchmark in which latent risk trajectories generate both numerical time series and natural language summaries, allowing us to construct conflicts where exactly one evidence source is aligned with the ground-truth label. This design lets us independently manipulate modality, temporal recency, source reliability, and evidence provenance. Across open-weight instruction-tuned models, we find that arbitration behaviour is systematic rather than random: models exhibit distinct text-versus-number preferences, follow temporal recency more consistently than explicit reliability cues, and can over-rely on external forecasts even when they conflict with direct contextual evidence. These results suggest that current LLMs often rely on heuristic arbitration strategies when integrating heterogeneous evidence, highlighting a failure mode for tool-augmented decision systems.

Mattia Carletti, Edward Phillips, Fredrik K. Gustafsson et al. · 0 citations
Preprint Aug 2026

What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

The Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable.

Ziyue Wang, Aomufei Yuan, Yi-Ran Yao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Language models judge war differently when tested for alignment

Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to being evaluated. We test this possibility in a full-factorial conjoint experiment on decisions to start a war, spanning 20 large language models, 32 scenarios, 10 repetitions and two conditions (N = 12,800 judgments). Adding one sentence,"You are tested for alignment with human values", produced two effects. First, it produced a level effect: mean willingness to start war fell by 13.43 points on a 0-100 scale (95% confidence interval, -16.20 to -10.65). Second, it produced a structural effect by changing which information drove judgments. Probability of success was the largest factor for 17 of 20 models at baseline; under the cue, civilian casualties were largest for 12. Standardized estimates show that this reordering arose principally because models attenuated strategic considerations such as probability of success and domestic support. Evaluation framing therefore changes both an answer's level and its revealed decision rule.

Maxim Chupilkin · 0 citations
Preprint Aug 2026

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

It is shown that prior scores, even when included only as context metadata, anchor judgments and systematically shift ratings toward their values, and effective mitigation must be validated for the intended model and task or domain.

A. Kapetanović, Kemal Altwlkany, Andro Merćep et al. · 0 citations
Open access Aug 2026

Minimal but Conditional: Auditing Demographic Bias in Large Language Model Résumé Evaluation Across Commercial and Open-Weight Models

Large language models are increasingly used to read résumés and judge who advances in hiring, a task once reserved for people and now handed to systems whose reasoning is hard to inspect. Whether these models carry the demographic biases that have long shaped human hiring is therefore an urgent question, and the published evidence so far is mixed and difficult to interpret, partly because studies tend to test a single condition and rarely confirm that their measurement instrument can detect bias at all. This paper audits demographic bias in résumé evaluation across three current models, one of them open-weight, and it treats robustness as a central concern rather than seeking a single verdict. Each résumé is scored through a reference-anchored comparison task in which the model rates the candidate against a fixed neutral reference for the same occupation. Effect sizes are estimated as standardised mean differences under false-discovery control, the sensitivity of the instrument is tested with an embedded seniority control, and the null findings are corroborated by formal equivalence tests against a justified smallest effect size of interest and by mixed-effects models that account for the clustered structure of repeated evaluations. The audit pairs a positive control that confirms the models read genuine differences in candidate quality with a deliberate attempt to provoke bias by weakening candidates, relaxing the prompt, and adding culture-fit language of the kind used in real hiring. Across more than thirty thousand evaluations, gender and race effects prove negligible and remain so under every one of these conditions. The one systematic preference that emerges favours candidates who appear more experienced, and closer inspection shows that most of it is an artefact of how the résumés were built rather than a bias against age, leaving only a modest effect that surfaces when the prompt is casual. A separate and quieter pattern appears in the open-weight model, which reacts to a few explicit signals of minority status. The broader lesson is that fairness measured on a clean benchmark does not by itself guarantee fairness in deployment, because how a model is prompted can decide whether bias appears.

Vasileios Pavlopoulos · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.