A tension between hallucination avoidance and user satisfaction is revealed and the importance of designing balanced refusal strategies is highlighted, to highlight the importance of designing balanced refusal strategies.
Abstract
While refusal-based safeguards to mitigate hallucinations in large language models (LLMs) are becoming increasingly common, they may conflict with users'preferences for definitive answers. However, we know little about how users respond to refusals across repeated interactions, when refusals become more or less acceptable, and for whom. In this work, we examine how refusal frequency, explanations, and need for cognitive closure (NFCC) shape responses to AI refusals. Participants (N=599) interacted with an AI system that never refused, refused infrequently, or refused frequently, with refusals either explained or unexplained. Participants were most satisfied with genuine responses, followed by hallucinations and then refusals, despite recognizing hallucinations as less accurate. Explanations increased satisfaction with infrequent, but not frequent, refusals. Higher-NFCC participants evaluated AI systems that refused more negatively. These findings reveal a tension between hallucination avoidance and user satisfaction and highlight the importance of designing balanced refusal strategies.
Many medical conditions require diagnosis through detailed, multi-context clinical assessment rather than from visual appearance alone. Despite this, vision-language models (VLMs) are increasingly queried to interpret images in ways that touch on medical or diagnostic judgments, raising safety concerns when such infere...
Karan Dua, Amit Agarwal, Hitesh Laxmichand Patel et al.· 0 citations
This work proposes the first taxonomy of LLM refusals that is grounded in pragmatic theory and finds that although models differ in how they refuse, their refusals are overall explicit and strongly morally evaluative, with interactional repair occurring mainly through offering or providing safer alternatives instead of...
Ruo-Xuan Li, Pin-Qiao Wang, Sheng Li et al.· 1 citation
People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-L...
Robin Welsch, Michelle Rausch, Pascal Knierim et al.· 0 citations
LLMs have become thinking companions for hundreds of millions worldwide, and they often do remarkable work. But their satisfaction-optimized design produces suppressed inquiry: the tendency for LLMs to inhibit rather than facilitate users' questioning process. LLMs appear fluent, confident, and seemingly all-knowing, t...
Z. Carmon, Itai Linzen, Aner Sela et al.· Current Opinion in Psycholog...· 0 citations
Users sometimes judge that an unfamiliar response is"so Claude."What does this judgment recognize, if it does not identify which model, process, conversation, or mind produced the response? I distinguish three orders of inquiry into AI identity. Constraint-first inquiry begins with conditions that a persisting interloc...
An understanding gap is illuminated: user attributions are partly guided by epistemic orientation and experiential ascriptions that make sophisticated simulation appear as understanding, especially in affective interaction, raising urgent questions about epistemic trust, relational vulnerability, and the ethics of AI c...
Erez Firt, Rinat B. Rosenberg-Kima· AI & SOCIETY· 0 citations
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduSep 30, 2026
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
MIT News · Artificial Intelligence· news.mit.eduSep 23, 2026
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.
AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.