Skip to content
Preprint

How Unlikely Is"Unlikely"? Assessing Verbal Probability Perception Across Large Language Models

Aug 2026 · 0 citations · 11 references
Computer Science

Abstract

Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty) remains unknown. We present a systematic cross-model evaluation using a word-to-number mapping task grounded in established human benchmarks. Eleven uncertainty expressions were presented to 19 models under two conditions, forced single-number response and explanation elicitation, alongside a novel bidirectional roundtrip test of internal consistency. LLMs track the human benchmark with surprising fidelity: word ordering is preserved, three anchor points are recovered, and ``possible''shows the highest variance and cross-model disagreement of any expression tested, consistent with its documented bimodal interpretation in humans. However, models show a systematic upward bias for negative expressions such as ``unlikely''and ``improbable.''Explanation elicitation reduces within-model variance while increasing between-model divergence, stabilizing individual models at the cost of inter-model consensus, and the roundtrip experiment reveals clear stratification, with frontier models maintaining coherent bidirectional representations. LLMs thus reproduce the structure of human verbal probability cognition, including its biases, while diverging systematically at the negative end---with implications for any setting where humans and models exchange probabilistic language.

View source

Similar papers

Jul 2026

The Computational Basis of Confidence in Large Language Models

A computational account of confidence in multimodal language models is provided, when answer logits behave as readouts of a latent decision variable is delineated, and statistical decision confidence is established as a unifying framework for studying confidence across biological and artificial intelligence.

D. Kumaran, Viorica Patraucean, M. Ovsjanikov et al. · 0 citations
Preprint Aug 2026

How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans

Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper examines whether LLMs can assess social attraction from theory-grounded persona profiles constructed from ten psychological and relational constructs and organized into three tiers: socially attractive, socially mixed, and socially unattractive. We examine LLM ratings in two studies and compare them with human judgments in a third study. In Study 1, 34 LLMs rated 12 profiles across three repeated runs. Although some models tended to give higher or lower ratings overall, they showed strong stability across runs, consistent three-tier ordering, and high agreement in relative profile ordering. Study 2 examined sensitivity to gender presentation using six matched name-and-pronoun profile pairs and a separate pronoun-only test with a gender-neutral name, finding no significant effects in either analysis. In Study 3, 198 human participants evaluated the six matched profiles from Study 2. Their ratings reproduced the three-tier structure and followed a profile ordering consistent with that of the LLMs. However, LLMs rated attractive profiles more positively and unattractive profiles more negatively than humans, while neither group showed a significant overall effect of gender presentation.

H. Mahmud, Khawaja Abaid Ullah, M. J. Khojasteh et al. · 0 citations
Preprint Aug 2026

Assessing mentalization in humans and large language models

Different capacities for mentalization across LLMs are demonstrated, and cognitive computational modeling is highlighted as a formal method for assessing comparative intelligence across humans and machines.

Aamir Sohail, Xintong Zhong, Arkady Konovalov et al. · 0 citations
#artificial intelligence Preprint Aug 2026

When Linguistic and Internal Confidence Diverge in Large Language Models

Regression analyses show that distributional properties of confidence scores explain much of the observed alignment pattern, with model metadata playing a smaller role after controls, and support a lossy-channel view of linguistic confidence.

Hefan Zhang, Bing-Quan Zhang, Ming Cheng et al. · 0 citations
Preprint Aug 2026

It's How You Ask: Gender-Associated Linguistic Bias in LLMs

Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses across three document types and four models. These effects persist after controlling for prompt complexity and feature carry-over. Explicit gender cues like sign-off names are encoded in the same representational space as linguistic dialect - suggesting shared underlying mechanisms - yet linguistic register is far more influential, producing large, consistent effects where names produce none. Our results further reveal that post-hoc mitigation is challenging: because these patterns are culturally embedded and outside conscious control, users cannot easily avoid them through strategic self-presentation, and mechanistic analysis reveals that linguistic features are encoded in early transformer layers and entangled with other features. Our work calls for upstream consideration of the influences of linguistic variation to mitigate disparate impacts of LLM-mediated workplace communication.

K. V. Koevering, Anjalie Field · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.