Metacognition—assessing the quality of one’s own cognitive performance—guides adaptive behaviour across species. Confidence signals can be extracted from language model outputs, yet a fundamental question remains: do models actually use these signals to decide whether to answer or abstain? Here we developed a four-phas...
D. Kumaran, N. Daw, Simon Osindero et al.· Nature Machine Intelligence· 11 citations
Behavioral foundation models have been proposed as stand-ins for human participants across settings, but it is unclear whether theories discovered on them generalize to humans or merely characterize the simulator. We ran the Automated Cognitive Scientist (\textsc{AutoCog}), a closed-loop discovery system in which LLM a...
A. Jagadish, Younes Strittmatter, Nori Jacoby et al.· 0 citations
A computational account of confidence in multimodal language models is provided, when answer logits behave as readouts of a latent decision variable is delineated, and statistical decision confidence is established as a unifying framework for studying confidence across biological and artificial intelligence.
D. Kumaran, Viorica Patraucean, M. Ovsjanikov et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.