Skip to content
Preprint

Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models

Aug 2026 · 0 citations · 20 references
Computer Science

TL;DR

The results show that LLMs systematically adapt their responses to align with prompt framing, even in factual contexts, which suggests that prompt framing can outweigh factual consistency in model responses.

Abstract

It is well established that large language models (LLMs) are sensitive to prompt framing, reflecting patterns in their training data or prior prompts. In this study, we investigate the extent to which LLMs reinforce users biases expressed in the prompts and examine the boundary between implicit framing effects and explicit prompt manipulation. Specifically, we evaluate how susceptible LLMs are to direct and suggestive prompts that encourage models to support or challenge particular positions. We evaluate six LLMs using 160 distinct prompts spanning ten topics across opinion-based and factual domains. The prompts systematically vary in prompting strategy, support versus challenge instructions, prompt polarity, users'expressed beliefs, and topic domain, spanning both opinion-based and factual questions. Our results show that LLMs systematically adapt their responses to align with prompt framing, even in factual contexts. This suggests that prompt framing can outweigh factual consistency in model responses. Overall, our findings delineate the extent and boundaries of LLM manipulability. Furthermore, the results imply that LLMs can reinforce subtle user biases and are susceptible to explicit prompt manipulation even in domains where responses should remain factually stable.

View source

Similar papers

Open access Jul 2026

How Does Prompt Anchoring Affect Large Language Model Outputs?

The study identifies prompt anchoring as a source of methodological variation in LLM-assisted content analysis, indicating that anchoring strategies should be explicitly specified, justified, and reported as part of the study methodology.

Eungi Kim · 0 citations
Preprint Aug 2026

It's How You Ask: Gender-Associated Linguistic Bias in LLMs

Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses across three document types and four models. These effects persist after controlling for prompt complexity and feature carry-over. Explicit gender cues like sign-off names are encoded in the same representational space as linguistic dialect - suggesting shared underlying mechanisms - yet linguistic register is far more influential, producing large, consistent effects where names produce none. Our results further reveal that post-hoc mitigation is challenging: because these patterns are culturally embedded and outside conscious control, users cannot easily avoid them through strategic self-presentation, and mechanistic analysis reveals that linguistic features are encoded in early transformer layers and entangled with other features. Our work calls for upstream consideration of the influences of linguistic variation to mitigate disparate impacts of LLM-mediated workplace communication.

K. V. Koevering, Anjalie Field · 0 citations
#natural language process... Preprint Sep 2026

The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis

Large Language Models (LLMs) increasingly use user context such as memory, profiles, and role prompts to personalize their responses. This personalization can affect evidence-based judgment: the same evidence may lead to different conclusions under different user contexts. Finance provides a high-stakes setting to study this problem because decisions often depend on interpreting long and complex documents. We test this using 3,575 SEC filings across twelve LLMs. We compare persona-conditioned retrieval, neutral retrieval, and memory-framed context to separate the effect of evidence selection from the effect of interpretation. We find that most user-context spillover comes from how models interpret the same evidence under different roles, rather than from retrieving different evidence. We then test two simple mitigation strategies: expressing the same investor mindset as a user profile instead of an assistant role, and separating evidence-based and personalized outputs. Both reduce spillover, but neither removes it completely, and their effectiveness varies substantially across models.

Ahmed Asaad, Amr Mohamed, Yang Zhang et al. · 0 citations
Open access 2026

When the Audit Depends on the Auditor: Prompt Sensitivity in Matched-Pair Audits of Large Language Models

Bias audits of large language models (LLMs) typically use a single prompt to compare model responses to identity-matched stimuli. This design rests on the implicit assumption that the estimated identity gap is stable across plausible prompt wordings. We test this assumption in a factorial experiment on three contemporary LLMs (Qwen, Llama, and GPT4o), which evaluated ten matched-text managerial vignettes under four prompts crossed with gender and name-signalled ethnicity manipulations, yielding 48,000 model calls. Prompt choice moved both rating levels and, in some cells, the estimated identity gap. Prompt effects on rating levels were often larger than identity effects, especially when a deliberately critical stress-test prompt was included. Because that critical prompt changes the evaluative construct rather than only its wording, we report results with and without it, treating the three construct-preserving prompts as the primary comparison and the critical prompt as an upperbound stress test. Where the gap was non-trivial, sign reversals were rare; the substantive importance of the prompt-induced instability we observed depends strongly on whether the underlying gap is itself non-trivial. We summarise cross-prompt variation with the Audit Disagreement Index (ADI). Single-prompt audits are not necessarily wrong, but they may be incomplete: reporting across multiple prompts allows readers to distinguish near-null gaps from substantively meaningful and stable ones.

Unknown authors · 0 citations
Review Aug 2026

Effects of Answer Format Variation on Gender Bias in Large Language Models

It is found that answer format does substantially alter measured outcomes, including reversals in order rankings, and the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment is highlighted.

K. Merzlyakova, Sebastian Padó, Franziska Weeber · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.