Current evidence supports clinician-supervised use of artificial intelligence (AI) systems rather than autonomous diagnosis, pending prospective and specialty-specific evaluation.
Abstract
Errors in differential diagnosis often arise while clinicians are generating and comparing candidate explanations. This review examines the use of large language models (LLMs) for this part of diagnostic reasoning. Internal medicine and pediatrics are the main focus; evidence from radiology, surgical subspecialties, infectious disease, and mental health is used to examine how findings change across specialties. Reported performance depends on the clinical setting, the quality of the input, the prompt, model adaptation, and the evaluation design. Some studies place LLMs near trainees and find that they produce wider, better-organized differentials. Experienced clinicians, however, remain more reliable overall. Domain adaptation, external knowledge, and interactive workflows have improved performance in specific evaluations, but hallucinations and automation bias remain, alongside unresolved questions of governance. Current evidence therefore supports clinician-supervised use of artificial intelligence (AI) systems rather than autonomous diagnosis, pending prospective and specialty-specific evaluation.
A narrative review of the available evidence presents a narrative review of the available evidence on the effect of LLMs on diagnostic reasoning, the optimal design of clinician-LLM interaction, the appropriate timing of consultation during the clinical encounter, the safest models of clinical-AI integration, and the main risks associated with their use.
L. Corral-Gudino, M. Ramos-Casals, M. Marcos et al.· Medicina clínica (Ed. impres...· 0 citations
: Large Language Models (LLMs) have shown potential to improve psychiatric assessment by addressing limitations in traditional diagnostic methods. This review summarizes recent developments in applying LLMs to mental health evaluation, focusing on three core strategies: questionnaire emulation (e.g., GPT-based PHQ-9/GAD-7), free-text classification of clinical narratives and social media posts, and integrated diagnostic-treatment workflows based on DSM/ICD criteria. Emphasis is placed on prompt engineering techniques — such as chain-of-thought and diagnostic reasoning prompts — that enhance model interpretability by generating stepwise rationales. These methods allow LLMs to mimic clinical reasoning while producing transparent, structured outputs. Empirical studies report high internal consistency and moderate-to-strong agreement with validated tools, along with performance metrics that approach or surpass human baselines in selected tasks. Key challenges include generalization across cultural contexts, explanation fidelity, and clinical applicability. Addressing these issues will require robust prompt design, alignment with clinical guidelines, and validation in real-world settings. This review provides a framework for understanding LLM-based diagnostic methods and outlines directions for future development in computational psychiatry.
Zhihao Li· Proceedings of the 3rd Inter...· 0 citations
Although promising, LLM-based systems are not yet reliable enough for autonomous medical diagnosis, and multiple recommendations for future research are contained to ensure a high level of safety, transparency, and clinical applicability for LLMs and other AI/ML-related technologies and devices.
M. U. K. Gunawardhna, Pirunthavi Wijikumar, D. Weerasinghe· Sri Lankan Journal of Applie...· 0 citations
A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.
Lalwani Saurabh, Bodetti Dr.Vishala, Gor Kishan et al.· Indian Journal of Computer S...· 0 citations
It is concluded that AI tools have matured to the extent that they can improve diagnostic decision-making of clinicians and can assist institutions in increasing diagnostic safety.
Findings show that clinical LLM explainability has shifted toward fluent generative rationales, but evidence that such explanations reflect model reasoning remains limited, and three regulatory priorities are highlighted: prioritizing explanations that enable independent verification or logic auditing over plausibility-only rationales; preferring inspectable models where regulatory documentation is required; and prospectively validating explanations in clinical workflows before scaling.