The results show that LLMs systematically adapt their responses to align with prompt framing, even in factual contexts, which suggests that prompt framing can outweigh factual consistency in model responses.
Abstract
It is well established that large language models (LLMs) are sensitive to prompt framing, reflecting patterns in their training data or prior prompts. In this study, we investigate the extent to which LLMs reinforce users biases expressed in the prompts and examine the boundary between implicit framing effects and explicit prompt manipulation. Specifically, we evaluate how susceptible LLMs are to direct and suggestive prompts that encourage models to support or challenge particular positions. We evaluate six LLMs using 160 distinct prompts spanning ten topics across opinion-based and factual domains. The prompts systematically vary in prompting strategy, support versus challenge instructions, prompt polarity, users'expressed beliefs, and topic domain, spanning both opinion-based and factual questions. Our results show that LLMs systematically adapt their responses to align with prompt framing, even in factual contexts. This suggests that prompt framing can outweigh factual consistency in model responses. Overall, our findings delineate the extent and boundaries of LLM manipulability. Furthermore, the results imply that LLMs can reinforce subtle user biases and are susceptible to explicit prompt manipulation even in domains where responses should remain factually stable.
The study identifies prompt anchoring as a source of methodological variation in LLM-assisted content analysis, indicating that anchoring strategies should be explicitly specified, justified, and reported as part of the study methodology.
Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses across three document types and four models. These effects persist after controlling for prompt complexity and feature carry-over. Explicit gender cues like sign-off names are encoded in the same representational space as linguistic dialect - suggesting shared underlying mechanisms - yet linguistic register is far more influential, producing large, consistent effects where names produce none. Our results further reveal that post-hoc mitigation is challenging: because these patterns are culturally embedded and outside conscious control, users cannot easily avoid them through strategic self-presentation, and mechanistic analysis reveals that linguistic features are encoded in early transformer layers and entangled with other features. Our work calls for upstream consideration of the influences of linguistic variation to mitigate disparate impacts of LLM-mediated workplace communication.
Large Language Models (LLMs) increasingly use user context such as memory, profiles, and role prompts to personalize their responses. This personalization can affect evidence-based judgment: the same evidence may lead to different conclusions under different user contexts. Finance provides a high-stakes setting to study this problem because decisions often depend on interpreting long and complex documents. We test this using 3,575 SEC filings across twelve LLMs. We compare persona-conditioned retrieval, neutral retrieval, and memory-framed context to separate the effect of evidence selection from the effect of interpretation. We find that most user-context spillover comes from how models interpret the same evidence under different roles, rather than from retrieving different evidence. We then test two simple mitigation strategies: expressing the same investor mindset as a user profile instead of an assistant role, and separating evidence-based and personalized outputs. Both reduce spillover, but neither removes it completely, and their effectiveness varies substantially across models.
Ahmed Asaad, Amr Mohamed, Yang Zhang et al.· 0 citations
Bias audits of large language models (LLMs) typically use a single prompt to compare
model responses to identity-matched stimuli. This design rests on the implicit assumption
that the estimated identity gap is stable across plausible prompt wordings. We test this
assumption in a factorial experiment on three contemporary LLMs (Qwen, Llama, and GPT4o), which evaluated ten matched-text managerial vignettes under four prompts crossed with
gender and name-signalled ethnicity manipulations, yielding 48,000 model calls. Prompt
choice moved both rating levels and, in some cells, the estimated identity gap. Prompt
effects on rating levels were often larger than identity effects, especially when a deliberately
critical stress-test prompt was included. Because that critical prompt changes the evaluative
construct rather than only its wording, we report results with and without it, treating the three
construct-preserving prompts as the primary comparison and the critical prompt as an upperbound stress test. Where the gap was non-trivial, sign reversals were rare; the substantive
importance of the prompt-induced instability we observed depends strongly on whether the
underlying gap is itself non-trivial. We summarise cross-prompt variation with the Audit
Disagreement Index (ADI). Single-prompt audits are not necessarily wrong, but they may be
incomplete: reporting across multiple prompts allows readers to distinguish near-null gaps
from substantively meaningful and stable ones.
Unknown authors· Advances in Artificial Intel...· 0 citations
It is found that answer format does substantially alter measured outcomes, including reversals in order rankings, and the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment is highlighted.
K. Merzlyakova, Sebastian Padó, Franziska Weeber· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.