Aug 2026· Frontiers in Medicine· Vol 13· 0 citations· 41 references
Medicine
TL;DR
The LLM showed offline potential for identifying cost-communication cues and generating clinician-facing preparation checklists for TKA, but content omissions, quality variation, safety risks, and substantive cross-run variability remained.
Abstract
Background and aims Costs associated with total knee arthroplasty (TKA) may affect treatment preparation, expectation management, and postoperative care planning. Previous large language model (LLM) studies have focused mainly on medical question answering, patient education, and clinical decision support, whereas their performance in clinician-facing preoperative cost-communication preparation remains unclear. This study evaluated an LLM using real-world clinical records from two hospitals in an expert-referenced offline evaluation. Methods Preoperative medical records of 80 patients who underwent primary unilateral TKA at two hospitals in China from January to May 2026 were included. A structured expert-panel process was used to develop a preoperative cost-communication framework comprising 4 dimensions and 17 clinical cues and to establish case-level minimum necessary communication items (CL-MNCIs) for each case. Task 1 assessed identification of the 17 cues. Task 2 assessed CL-MNCI coverage and classified all generated items according to case relevance, medical-record support, redundancy, and safety. Twenty-four cases were non-randomly selected by CL-MNCI count for three repeated-generation runs. Results Task 1 comprised 1,360 case–label classification units. Micro-precision, micro-recall, micro-F1, accuracy, MCC, and macro-F1 were 0.901, 0.888, 0.894, 0.948, 0.860, and 0.840, respectively. Experts established 569 CL-MNCIs, of which 483 were covered, yielding an overall coverage rate of 84.9%; complete coverage was achieved in 17 cases. The LLM generated 716 items, including 12 safety events across 9 cases, 483 items matching CL-MNCIs, 132 record-supported supplementary items, 50 redundant items, and 39 items with insufficient record support. Overall, 665 items (92.9%) were case-relevant and record-supported, although this proportion included redundant content. Pairwise Jaccard similarity for covered CL-MNCI sets ranged from 0.813 to 0.823. Of 180 CL-MNCIs, 126 (70.0%) were covered in all three runs, and case-level agreement in safety classification ranged from 87.5 to 95.8%. Conclusion The LLM showed offline potential for identifying cost-communication cues and generating clinician-facing preparation checklists for TKA, but content omissions, quality variation, safety risks, and substantive cross-run variability remained. Its use should be limited to clinician-reviewed communication preparation and should not replace direct patient cost disclosure or professional judgment.
BACKGROUND
Large language models (LLMs) may help organize clinical information, but their use in perioperative settings requires careful evaluation because errors may have immediate safety implications. This study aimed to describe the feasibility and perceived usefulness of a single-centre, expert-corrected LLM workflow for preparing preoperative anesthesia assessment drafts for complex consultation cases. Secondary aims were to describe error patterns identified by anesthesiologists and to explore residents' perceptions after reviewing expert-corrected materials.
METHODS
We retrospectively selected 15 complex preoperative anesthesia consultation cases from The First Affiliated Hospital, Zhejiang University School of Medicine. Complex cases were defined as cases referred for specialized preoperative anesthesia consultation because of multiple comorbidities, clinically relevant organ dysfunction, implanted devices, major cardiopulmonary or vascular disease, or other features requiring individualized anesthetic assessment. A large language model-based artificial intelligence system, DeepSeek-R1, was used to generate structured draft assessments from de-identified case information and a standardized prompt. The original LLM-generated drafts were independently reviewed by three experienced anesthesiologists for information completeness, scientific plausibility, and focus on key risk factors. All resident-facing documents were then corrected by senior anesthesiologists before being distributed to first-year anesthesia residents, who completed a questionnaire on perceived usefulness, guidance value, educational support, and error recognition. No comparator group, blinding, or objective performance outcome was included.
RESULTS
The LLM-generated drafts were structurally complete but contained clinically relevant errors, mainly in risk assessment and anesthetic plan formulation. Thirty-two error instances were identified across the 15 cases, including errors related to proposed plans, American Society of Anesthesiologists (ASA) physical status classification, and abnormal test interpretation. After expert correction, residents reported favorable perceptions of the materials, but their self-reported ability to identify expert-annotated errors remained limited. These findings should be interpreted as perception-based observations rather than evidence of educational effectiveness.
CONCLUSIONS
In this single-centre exploratory study, an LLM-assisted workflow was feasible for producing preliminary preoperative anesthesia assessment drafts, but expert correction was essential before resident-facing use. The findings do not establish standalone clinical or educational effectiveness of the LLM. Instead, they highlight the gap between structural completeness and clinical reliability, and suggest that any use of LLM-generated anesthesia assessment drafts should remain adjunctive, supervised, and explicitly framed to reduce automation bias.
Shuhan Gu, Jian-Fan Ping, Hui-dan Lin et al.· Medicina clínica (Ed. impres...· 0 citations
INTRODUCTION
Large language models (LLMs) are increasingly used as sources of information across many fields, including healthcare. As patients turn to these models for health-related queries, evaluating the accuracy and reliability of their responses is essential. Peri-acetabular osteotomy (PAO) is performed on younger patients - a group more likely to use digital tools like LLMs for health information. This study assesses the accuracy and readability performance of two leading LLMs, ChatGPT and Google Gemini, in addressing common patient questions on PAO.
METHODS
A panel of fellowship-trained PAO surgeons created ten commonly asked patient questions based on real-world experience. Responses from each LLM were assessed by the same three surgeons, blinded to response origin, using a 5-point Likert scale to evaluate clarity, accuracy, and completeness. Readability was measured with Flesch-Kincaid Reading scores.
RESULTS
ChatGPT outperformed Gemini with an average score of 4.17 vs 3.13 (t = -3.08, p = 0.006). ChatGPT's responses were often rated higher for completeness and clarity, particularly in areas needing detailed explanation, and usually required minimal clarification. Gemini sometimes lacked specificity or included minor inaccuracies that reduced its perceived reliability. Both LLMs produced responses with similar "difficult" Flesch Reading Ease scores.
CONCLUSIONS
There may be significant differences in how effectively LLMs support patients with surgical queries. ChatGPT more consistently met expert standards for clarity and thoroughness. As LLM usage expands, ChatGPT may aid patient education on hip surgery, supporting consultations, informed decisions and postoperative guidance.
TP Davis, B. Guevel, K. Logishetty et al.· Annals of the Royal College...· 0 citations
BACKGROUND
Robotic-assisted total knee arthroplasty (rTKA) is increasingly used because of its surgical precision. However, inconsistent outcomes and high costs often lead patients to seek additional information from artificial intelligence (AI) tools. Large language models (LLMs) such as ChatGPT-4o, Gemini-2.5-Flash, and DeepSeek-V3 are commonly used, but their reliability and readability in orthopaedics remain unclear.
OBJECTIVES
To compare the reliability, usefulness, quality, and readability of responses to common patient questions about rTKA generated by leading LLMs.
METHODS
Three LLMs answered 20 frequently asked patient questions (n = 20) identified through Google Trends and expert validation. Three orthopaedic specialists (n = 3) evaluated reliability, usefulness, and overall quality using validated scales, while readability was assessed with standard indices.
RESULTS
Inter-rater reliability was good to excellent (ICC = 0.728-0.879). Gemini-2.5-Flash achieved significantly higher reliability and usefulness scores than ChatGPT-4o and DeepSeek-V3 (all p < 0.05). ChatGPT-4o and DeepSeek-V3 produced more readable but less accurate content, revealing an inverse relationship between reliability and readability.
CONCLUSIONS
Gemini-2.5-Flash provided the most reliable responses, highlighting the need for supervised integration of LLMs in patient education.
Mehmet Utku Çiftçi, A. Koluman, Ebru Aloğlu Çiftçi et al.· Knee (Oxford)· 0 citations
LLMs can support communication and education following spine surgery when used with structured prompting when used with structured prompting and ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers.
S. Wegmann, T. Rosenkranz, Philipp Egenolf et al.· European spine journal· 0 citations