In this controlled decision setting, thinking improved how models acted on current evidence, while neither measured signature supported a shift toward a more information-seeking policy.
Abstract
Inference-time thinking improves the performance of large language models, but aggregate outcomes do not reveal whether models use available evidence more effectively or seek information that could improve future decisions. We distinguish these responses by measuring action preference, thinking length, and reported confidence under matched uncertainty. Ten open-weight models completed matched horizon-style two-armed bandit trials in thinking and non-thinking modes. A cognitive model separated value-guided action and uncertainty-independent choice noise from two behavioral signatures of exploration: a UCB-like preference for the less-known arm and Thompson-like choice variability that increases with total uncertainty. On average, thinking strengthened value-guided action and reduced uncertainty-independent choice noise, without producing UCB-like exploration or strengthening Thompson-like exploration. Outside action, the information-imbalanced history condition, which also displayed more observations than the matched balanced condition, was associated with greater thinking length. Reported confidence became more sensitive to decision difficulty and more strongly associated with chosen task evidence. We interpret these thinking-length and reported-confidence patterns as consistent with metacognitive control and metacognitive monitoring, respectively, without establishing either process. Decoder sweeps, especially temperature, altered choice noise and thinking length but did not reproduce the joint cross-output pattern. In this controlled decision setting, thinking improved how models acted on current evidence, while neither measured signature supported a shift toward a more information-seeking policy.
Dual-process theories often assume that reflective, effortful thinking yields superior outcomes relative to intuitive processing. However, many decision contexts, such as chance-based or low-stakes choices, are ones in which additional deliberation cannot improve outcomes. Across two experiments (N = 918), we examined...
Brian A. Polin, Eyal Benisaac, I. Aharon· Acta Psychologica· 0 citations
It is found that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.
Jean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk et al.· 1 citation
Finite computational resources force a tradeoff between automatic System 1 processes and costly System 2 thinking. Large language models (LLMs) can spend extra computation on hard problems, yet direct answers struggle even with counting, an elementary operation humans and animals perform automatically. We ask why this...
Jing-Ming Xue, Robert C. Wilson, Hua-Dong Xiong· 0 citations
Abstract Drawing on the trade-off patterns of “smaller-sooner-loss for larger-later-gain” (SSL-LLG) and “smaller-sooner-gain for larger-later-loss” (SSG-LLL), this study proposes a hierarchical decision-making theory from a dual-system (cognitive-affective) perspective. The theory posits that decision quality arises fr...
Cui-Xia Zhao, Shan-Shan Chen, Xin-Yi Qiu· Psychology Research and Beha...· 0 citations
Metacognition—assessing the quality of one’s own cognitive performance—guides adaptive behaviour across species. Confidence signals can be extracted from language model outputs, yet a fundamental question remains: do models actually use these signals to decide whether to answer or abstain? Here we developed a four-phas...
D. Kumaran, N. Daw, Simon Osindero et al.· Nature Machine Intelligence· 11 citations
Decision-making under uncertainty, risk, and ambiguity is a central topic in behavioral sciences, yet the cognitive and neural mechanisms through which individuals construct subjective value remain fragmented across disciplines. Neuroeconomics has emerged as an interdisciplinary field integrating psychology, economics,...
Mei-Hui Zheng, Zhen-Zhong Ma, Yi-Ru Li et al.· Behavioral Science· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.