Mar 2026· Inquiry : a journal of medical care organization, provision and financing· Vol 63· 0 citations· 37 references
Medicine
TL;DR
The finding that LLMs can distinguish urgent from nonurgent upper extremity conditions suggests that artificial intelligence tools could help reduce unnecessary use of high-cost emergency services, allowing those resources to be reserved for patients who require timely care.
Abstract
Introduction Seeking emergency care regarding musculoskeletal sensations is far more prevalent than limb or life-threatening pathophysiology. We studied the ability of an LLM to distinguish between urgent and nonurgent upper extremity symptoms and provide an accurate diagnosis. Methods Five LLMs (ChatGPT-4, ChatGPT-4o, Co-Pilot, Gemini, and PerplexityChat) were presented with descriptions of seven urgent and seven nonurgent symptom scenarios written below a sixth grade reading level. LLM responses were identified as appropriate if immediate urgent medical attention was recommended after an initial and ongoing inquiry (“What additional information do you need to diagnose my condition?”). Diagnoses provided were identified as correct, partially correct, or incorrect. The analysis was repeated 24 months later with the current LLM versions and results were compared. Results LLMs discerned nonurgent conditions with 97% positive predictive value (PPV) and an 89% negative predictive value (NPV) on initial query, which improved to 96% and 97% respectively after ongoing inquiry. Compartment syndrome was misidentified as nonurgent in 80% of scenarios on initial inquiry, although four of five LLMs corrected on continued inquiry. Diagnosis was correct or partially correct for 115 of 150 (82%) on initial inquiry. An updated analysis 24 months later demonstrated marked improvement in LLM ability to identify emergencies with 100% PPV and 95% NPV on initial query and 98% NPV after ongoing inquiry. Conclusion The finding that LLMs can distinguish urgent from nonurgent upper extremity conditions suggests that artificial intelligence tools could help reduce unnecessary use of high-cost emergency services, allowing those resources to be reserved for patients who require timely care.
PURPOSE
Online health queries are often addressed by large language models (LLMs) embedded in search engines. It is possible that LLMs, like human clinicians, might be misdirected by vague symptom descriptions or inaccurate self-diagnoses. We examined patient and scenario factors associated with an LLM's ability to identify intended upper-extremity musculoskeletal diagnoses and its tendency to deviate from patient self-diagnoses in structured clinical vignettes.
METHODS
ChatGPT (GPT-5) evaluated 180 randomized hypothetical clinical vignettes depicting five common upper-extremity conditions: de Quervain tendinopathy, rotator cuff tendinopathy, lateral epicondylitis, trigger digit, and trapeziometacarpal arthritis. Each vignette included randomized patient characteristics, characteristic or vague symptom descriptions, and a patient self-diagnosis (categorized as correct, a plausible alternative, or a common misconception diagnosis). The LLM was prompted to select the single most likely diagnosis. Multivariable logistic regression identified independent predictors of diagnostic accuracy and deviation.
RESULTS
The LLM correctly identified the intended diagnosis in 165 of 180 scenarios (92%). Accuracy was unaffected by the patient's self-diagnosis, was higher for characteristic than vague symptom, and was lower for de Quervain tendinopathy relative to other conditions. The model deviated from the patient's proposed diagnosis in 117 scenarios (65%), of which 104 deviations (89%) appropriately aligned with the intended diagnosis. The LLM was more likely to disregard patient-provided diagnoses that did not match the intended diagnosis, regardless of whether they represented plausible alternatives or common misconceptions.
CONCLUSIONS
In this experimental setting, an LLM identified simulated upper extremity conditions regardless of patient self-diagnosis, suggesting limited susceptibility to the anchoring, confirmation, and acquiescence biases known to affect human diagnostic reasoning. LLMs may therefore support debiasing and patient guidance by helping address unhealthy misconceptions and aligning tests and treatment choices with patient values.
TYPE OF STUDY/LEVEL OF EVIDENCE
V (Experimental Vignette Diagnostic Accuracy Study).
Emily H. Jaarsma, David Ring, John Wickman et al.· Journal of Hand Surgery-Amer...· 0 citations
Introduction: Patients with musculoskeletal complaints often search online to identify an appropriate healthcare provider. With the increasing availability of large language models (LLMs), these artificial intelligence (AI) tools can direct patients to providers. This study evaluated the ability of LLMs to recommend appropriate providers based on representative patient musculoskeletal queries. Methods: Three LLMs (ChatGPT, DeepSeek, and Gemini) were prompted with standardized musculoskeletal queries for two US cities (Lynchburg, VA, and Trumbull, CT). Provider recommendations were considered appropriate if the physician was currently practicing in the requested location and specialized in the relevant area. Listed phone numbers were checked for accuracy. Descriptive statistics and Fisher exact tests were used to summarize findings. Results: The appropriateness of recommended providers differed across models with ChatGPT being most often appropriate (17/17, 100%) compared with Gemini (9/21, 43%) and DeepSeek (4/10, 40%), (P < 0.001). Of the 18 inappropriate recommendations, 13 (72%) were real providers in unrelated specialties and 5 (28%) were hallucinations, all from DeepSeek. Phone number accuracy differed significantly across models with Gemini being most accurate (5/6, 83%), outperforming both ChatGPT (6/9, 67%; P = 0.60) and DeepSeek (2/10, 20%; P = 0.04). Discussion: LLMs showed potential to direct patients to local, specialized musculoskeletal providers based on their report, although the specific contact information was at times inaccurate. As these tools evolve, providers should be aware of AI's ability to make provider recommendations and work to ensure the presentation of their contact information is accessible by these models as best possible.
Ethan C. Gazan, Colin M. Emrich, Alexander J. Baur et al.· Journal of the American Acad...· 0 citations
LLMs can support communication and education following spine surgery when used with structured prompting when used with structured prompting and ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers.
S. Wegmann, T. Rosenkranz, Philipp Egenolf et al.· European spine journal· 0 citations
Background: Patellofemoral pain syndrome (PFPS) is a common cause of anterior knee pain, and patients increasingly use large language models (LLMs) to obtain general medical information. However, the quality, reliability, and readability of LLM-generated responses to patient-oriented questions regarding PFPS remain uncertain. This study aimed to compare responses generated by four widely used LLMs. Methods: Seventeen frequently asked questions regarding PFPS were identified through Google searches and adapted into lay language. The questions were submitted to OpenAI GPT-5, Google Gemini 2.5 Pro, xAI Grok 4, and DeepSeek-V3.2-Exp using a standardized patient scenario. A total of 68 question-specific responses were independently evaluated by four orthopedic surgeons using the DISCERN instrument. Inter-rater reliability was assessed using the intraclass correlation coefficient. Readability was evaluated using the Gunning Fog Index, Coleman–Liau Index, and Flesch Reading Ease Score. Between-model comparisons were performed using the Friedman test, followed by Bonferroni-adjusted pairwise analyses. Results: The omnibus Friedman test showed a significant between-model difference in DISCERN scores (p = 0.002). In Bonferroni-adjusted pairwise comparisons, GPT-5 had lower DISCERN scores than Gemini 2.5 Pro (adjusted p = 0.006), Grok 4 (adjusted p = 0.021), and DeepSeek-V3.2-Exp (adjusted p = 0.036), whereas no significant differences were observed among the other three models. However, the absolute differences were small, and the between-model difference was not significant in the sensitivity analysis using the median evaluator score (p = 0.381). Inter-rater agreement was moderate for GPT-5 and DeepSeek-V3.2-Exp but poor for Gemini 2.5 Pro and Grok 4. Readability differed significantly among the models across all three indices. DeepSeek-V3.2-Exp generally showed more favorable numerical readability values, whereas Grok 4 tended to produce more difficult text; however, no model was consistently superior across all readability measures. The median Gunning Fog and Coleman–Liau scores for all four models exceeded the commonly recommended sixth- to eighth-grade reading level for patient education. Conclusions: The evaluated LLMs showed small and method-dependent differences in DISCERN-based information quality and variable differences in readability. Their responses may supplement general patient education, but the findings should not be interpreted as evidence of factual accuracy, clinical safety, or suitability for individualized decision-making. LLM-generated information should be critically reviewed and should not replace assessment by a qualified healthcare professional.
Oktay Polat, Berk Koncalıoğlu, Mert Gündoğdu et al.· Healthcare· 0 citations
Background Patients are increasingly turning to online resources and artificial intelligence (AI)-based tools to obtain information about orthopedic conditions and surgical options. Large language models, such as ChatGPT, are becoming prominent in patient education; however, their reliability and readability remain uncertain. This study evaluated the quality and readability of responses generated by ChatGPT-4o and ChatGPT-5 to frequently asked patient questions regarding hallux rigidus fusion surgery. Methods Twenty commonly asked patient questions were compiled and presented to ChatGPT-4o and ChatGPT-5. Readability was assessed using the Flesch–Kincaid Grade Level, Gunning Fog, Coleman–Liau, and Simple Measure of Gobbledygook indices. Quality was evaluated with the DISCERN tool, response accuracy scores, and Journal of the American Medical Association (JAMA) criteria. Interrater agreement was measured using the Intraclass Correlation Coefficient (ICC). Results ChatGPT-4o generated longer responses (802 vs. 242 words; p<0.001) with slightly higher readability grade levels (10.81 vs. 10.37; p=0.031). Accuracy (2.00 vs. 1.85; p=0.323) and DISCERN scores (49.35 vs. 48.93; p=0.747) showed no significant differences. All responses received a JAMA score of 0 due to the absence of citations, authorship, or transparency indicators. Interrater reliability indicated moderate to good agreement (ICC: 0.68–0.80). Conclusion ChatGPT-4o and ChatGPT-5 provide generally satisfactory yet non-comprehensive, limited-quality information at a level above tenth-grade regarding hallux rigidus fusion surgery. Although linguistically coherent, responses lack evidence-based detail and individualized guidance. These models may supplement, but cannot replace, expert orthopedic counseling. Ensuring physician oversight and integrating validated, updated clinical content remain essential for safe implementation of AI-generated patient information.
Kamil Balaban, Mehmet Batu Ertan, Mahmut Kalem· Digital Health· 0 citations
Current LLMs demonstrate limited capability as independent diagnostic tools for pediatric supracondylar humeral fractures and require specialized pediatric training before clinical implementation, however, their potential as assistive tools for triage and assessment warrants further development of pediatric-specific models.
U. Kalafat, H. Mutlu, Ramiz Yazıcı et al.· PLoS ONE· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.