Probing can be a cost-effective triage mechanism for routing LLM answers to human review and quality control procedures in high-stakes financial applications and shows that among confident answers, those for which all eight resamples agree, 15-23% are wrong on FinQA.
Abstract
Large language models (LLMs) in financial applications fail most consequentially when they are confidently wrong. Hedged, uncertain answers invite scrutiny, whereas confident errors silently degrade downstream decisions without warning. We ask how reliably such confidently wrong answers, or confident hallucinations, can be detected from a model's internal activations, and whether those activations carry information beyond its observable outputs. We train linear probes on the residual stream and evaluate them on two established question-answering (QA) benchmarks built from real filings, FinQA and TAT-QA. Behavioral confidence is measured as the agreement among eight resampled answers to the same question, and probe effectiveness is compared against baselines, such as token log-probabilities and the model's own True/False self-assessment of its answer. Our findings show that among confident answers, those for which all eight resamples agree, 15-23% are wrong on FinQA. There the probes have a significant advantage over baseline methods in detecting hallucinations, holding 0.68-0.77 AUROC while the best baselines fall to 0.55-0.63, across Qwen3-8B, Llama-3.1-8B, and Gemma-2-9B. Our results suggest that probing can be a cost-effective triage mechanism for routing LLM answers to human review and quality control procedures in high-stakes financial applications.
Large language models can hallucinate even when the knowledge required for a correct answer is already available. We study this failure through a latent-key view of inference, where answer selection depends on competition among associations acquired during pretraining. We show that model predictions can be highly sensi...
Xu-Han Tong, Hao-Yue Bai, Da-Wei Zhou et al.· 0 citations
Large Language Models (LLMs) are increasingly deployed in high-stakes domains such as legal assistance and consumer rights support, where response accuracy is of paramount importance; however, a critical challenge in such deployments is hallucination-the generation of factually incorrect or unverified statements that m...
Adithyan K, K. V., D. G.· International Conference on...· 0 citations
Hallucinations—fluent outputs containing incorrect or unsupported factual claims—remain an important obstacle to reliable use of large language models (LLMs). This study evaluates the scope and limits of HalluDetector, a reference-based detector that identifies contradiction-like evidence using lexical, numeric, unavai...
Seung-Ho Lee· Journal of high school scien...· 0 citations
ADAM-Bench (Auditing Dialogue Assertions with Multimodal Evidence), a benchmark for paper-grounded hallucinations in scientific dialogue, is introduced and two tasks are defined: hallucination detection and minimal evidence set localization.
Ze-Xing Zhang, Tian-Yang Lei, Ke-Wei Yang et al.· Proceedings of the 32nd ACM...· 0 citations
The lower option-selection accuracy indicates that distinguishing the verified answer from plausible alternatives remains more challenging than detecting hallucinated responses.
This work evaluates seven confidence estimators, three inference-only and four trained internal probes, across five open-weight LVLMs and four conditions from three financial visual question-answering benchmarks, one bilingual; every probe is trained only on natural images and applied to finance without adaptation, so...
Mohammad M. Ghassemi, Simerjot Kaur, Charese H. Smiley et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.