Skip to content

Thinking Under Uncertainty: Evidence Use and Information-Seeking in Language Models

Jul 2026 · arXiv.org · Vol abs/2607.26845 · 0 citations · 16 references
Computer Science

TL;DR

In this controlled decision setting, thinking improved how models acted on current evidence, while neither measured signature supported a shift toward a more information-seeking policy.

Abstract

Inference-time thinking improves the performance of large language models, but aggregate outcomes do not reveal whether models use available evidence more effectively or seek information that could improve future decisions. We distinguish these responses by measuring action preference, thinking length, and reported confidence under matched uncertainty. Ten open-weight models completed matched horizon-style two-armed bandit trials in thinking and non-thinking modes. A cognitive model separated value-guided action and uncertainty-independent choice noise from two behavioral signatures of exploration: a UCB-like preference for the less-known arm and Thompson-like choice variability that increases with total uncertainty. On average, thinking strengthened value-guided action and reduced uncertainty-independent choice noise, without producing UCB-like exploration or strengthening Thompson-like exploration. Outside action, the information-imbalanced history condition, which also displayed more observations than the matched balanced condition, was associated with greater thinking length. Reported confidence became more sensitive to decision difficulty and more strongly associated with chosen task evidence. We interpret these thinking-length and reported-confidence patterns as consistent with metacognitive control and metacognitive monitoring, respectively, without establishing either process. Decoder sweeps, especially temperature, altered choice noise and thinking length but did not reproduce the joint cross-output pattern. In this controlled decision setting, thinking improved how models acted on current evidence, while neither measured signature supported a shift toward a more information-seeking policy.

View source

Similar papers

Open access Sep 2026

When deliberation becomes inefficient.

Dual-process theories often assume that reflective, effortful thinking yields superior outcomes relative to intuitive processing. However, many decision contexts, such as chance-based or low-stakes choices, are ones in which additional deliberation cannot improve outcomes. Across two experiments (N = 918), we examined...

Brian A. Polin, Eyal Benisaac, I. Aharon · 0 citations
Preprint Aug 2026

Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

It is found that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.

Jean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk et al. · 1 citation
#machine learning Preprint Sep 2026

Counting on Thinking: Tracing Evidence Integration in Language Models

Finite computational resources force a tradeoff between automatic System 1 processes and costly System 2 thinking. Large language models (LLMs) can spend extra computation on hard problems, yet direct answers struggle even with counting, an elementary operation humans and animals perform automatically. We ask why this...

Jing-Ming Xue, Robert C. Wilson, Hua-Dong Xiong · 0 citations
Review Open access Sep 2026

Hierarchical Decision-Making Theory: A Cognitive-Affective Framework for Explaining Decision Quality

Abstract Drawing on the trade-off patterns of “smaller-sooner-loss for larger-later-gain” (SSL-LLG) and “smaller-sooner-gain for larger-later-loss” (SSG-LLL), this study proposes a hierarchical decision-making theory from a dual-system (cognitive-affective) perspective. The theory posits that decision quality arises fr...

Cui-Xia Zhao, Shan-Shan Chen, Xin-Yi Qiu · 0 citations
#machine learning Open access Mar 2026

Causal evidence that language models use confidence to drive behaviour

Metacognition—assessing the quality of one’s own cognitive performance—guides adaptive behaviour across species. Confidence signals can be extracted from language model outputs, yet a fundamental question remains: do models actually use these signals to decide whether to answer or abstain? Here we developed a four-phas...

D. Kumaran, N. Daw, Simon Osindero et al. · 11 citations
Review Open access Aug 2026

Subjective Values in Decision-Making Under Uncertainty: A Systematic Review

Decision-making under uncertainty, risk, and ambiguity is a central topic in behavioral sciences, yet the cognitive and neural mechanisms through which individuals construct subjective value remain fragmented across disciplines. Neuroeconomics has emerged as an interdisciplinary field integrating psychology, economics,...

Mei-Hui Zheng, Zhen-Zhong Ma, Yi-Ru Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.