Skip to content
Review Open access

Evaluation of a large language model for clinician-facing preoperative cost-communication preparation in total knee arthroplasty

Aug 2026 · Frontiers in Medicine · Vol 13 · 0 citations · 41 references
Medicine

TL;DR

The LLM showed offline potential for identifying cost-communication cues and generating clinician-facing preparation checklists for TKA, but content omissions, quality variation, safety risks, and substantive cross-run variability remained.

Abstract

Background and aims Costs associated with total knee arthroplasty (TKA) may affect treatment preparation, expectation management, and postoperative care planning. Previous large language model (LLM) studies have focused mainly on medical question answering, patient education, and clinical decision support, whereas their performance in clinician-facing preoperative cost-communication preparation remains unclear. This study evaluated an LLM using real-world clinical records from two hospitals in an expert-referenced offline evaluation. Methods Preoperative medical records of 80 patients who underwent primary unilateral TKA at two hospitals in China from January to May 2026 were included. A structured expert-panel process was used to develop a preoperative cost-communication framework comprising 4 dimensions and 17 clinical cues and to establish case-level minimum necessary communication items (CL-MNCIs) for each case. Task 1 assessed identification of the 17 cues. Task 2 assessed CL-MNCI coverage and classified all generated items according to case relevance, medical-record support, redundancy, and safety. Twenty-four cases were non-randomly selected by CL-MNCI count for three repeated-generation runs. Results Task 1 comprised 1,360 case–label classification units. Micro-precision, micro-recall, micro-F1, accuracy, MCC, and macro-F1 were 0.901, 0.888, 0.894, 0.948, 0.860, and 0.840, respectively. Experts established 569 CL-MNCIs, of which 483 were covered, yielding an overall coverage rate of 84.9%; complete coverage was achieved in 17 cases. The LLM generated 716 items, including 12 safety events across 9 cases, 483 items matching CL-MNCIs, 132 record-supported supplementary items, 50 redundant items, and 39 items with insufficient record support. Overall, 665 items (92.9%) were case-relevant and record-supported, although this proportion included redundant content. Pairwise Jaccard similarity for covered CL-MNCI sets ranged from 0.813 to 0.823. Of 180 CL-MNCIs, 126 (70.0%) were covered in all three runs, and case-level agreement in safety classification ranged from 87.5 to 95.8%. Conclusion The LLM showed offline potential for identifying cost-communication cues and generating clinician-facing preparation checklists for TKA, but content omissions, quality variation, safety risks, and substantive cross-run variability remained. Its use should be limited to clinician-reviewed communication preparation and should not replace direct patient cost disclosure or professional judgment.

Read PDF

Similar papers

Review Open access Aug 2026

A large language models-assisted and expert-corrected workflow for preoperative anesthesia assessment drafts: A single-centre exploratory feasibility study.

BACKGROUND Large language models (LLMs) may help organize clinical information, but their use in perioperative settings requires careful evaluation because errors may have immediate safety implications. This study aimed to describe the feasibility and perceived usefulness of a single-centre, expert-corrected LLM workflow for preparing preoperative anesthesia assessment drafts for complex consultation cases. Secondary aims were to describe error patterns identified by anesthesiologists and to explore residents' perceptions after reviewing expert-corrected materials. METHODS We retrospectively selected 15 complex preoperative anesthesia consultation cases from The First Affiliated Hospital, Zhejiang University School of Medicine. Complex cases were defined as cases referred for specialized preoperative anesthesia consultation because of multiple comorbidities, clinically relevant organ dysfunction, implanted devices, major cardiopulmonary or vascular disease, or other features requiring individualized anesthetic assessment. A large language model-based artificial intelligence system, DeepSeek-R1, was used to generate structured draft assessments from de-identified case information and a standardized prompt. The original LLM-generated drafts were independently reviewed by three experienced anesthesiologists for information completeness, scientific plausibility, and focus on key risk factors. All resident-facing documents were then corrected by senior anesthesiologists before being distributed to first-year anesthesia residents, who completed a questionnaire on perceived usefulness, guidance value, educational support, and error recognition. No comparator group, blinding, or objective performance outcome was included. RESULTS The LLM-generated drafts were structurally complete but contained clinically relevant errors, mainly in risk assessment and anesthetic plan formulation. Thirty-two error instances were identified across the 15 cases, including errors related to proposed plans, American Society of Anesthesiologists (ASA) physical status classification, and abnormal test interpretation. After expert correction, residents reported favorable perceptions of the materials, but their self-reported ability to identify expert-annotated errors remained limited. These findings should be interpreted as perception-based observations rather than evidence of educational effectiveness. CONCLUSIONS In this single-centre exploratory study, an LLM-assisted workflow was feasible for producing preliminary preoperative anesthesia assessment drafts, but expert correction was essential before resident-facing use. The findings do not establish standalone clinical or educational effectiveness of the LLM. Instead, they highlight the gap between structural completeness and clinical reliability, and suggest that any use of LLM-generated anesthesia assessment drafts should remain adjunctive, supervised, and explicitly framed to reduce automation bias.

Shuhan Gu, Jian-Fan Ping, Hui-dan Lin et al. · 0 citations
Open access Aug 2026

Evaluating large language models in patient education: a comparative analysis addressing frequently asked questions in peri-acetabular osteotomy.

INTRODUCTION Large language models (LLMs) are increasingly used as sources of information across many fields, including healthcare. As patients turn to these models for health-related queries, evaluating the accuracy and reliability of their responses is essential. Peri-acetabular osteotomy (PAO) is performed on younger patients - a group more likely to use digital tools like LLMs for health information. This study assesses the accuracy and readability performance of two leading LLMs, ChatGPT and Google Gemini, in addressing common patient questions on PAO. METHODS A panel of fellowship-trained PAO surgeons created ten commonly asked patient questions based on real-world experience. Responses from each LLM were assessed by the same three surgeons, blinded to response origin, using a 5-point Likert scale to evaluate clarity, accuracy, and completeness. Readability was measured with Flesch-Kincaid Reading scores. RESULTS ChatGPT outperformed Gemini with an average score of 4.17 vs 3.13 (t = -3.08, p = 0.006). ChatGPT's responses were often rated higher for completeness and clarity, particularly in areas needing detailed explanation, and usually required minimal clarification. Gemini sometimes lacked specificity or included minor inaccuracies that reduced its perceived reliability. Both LLMs produced responses with similar "difficult" Flesch Reading Ease scores. CONCLUSIONS There may be significant differences in how effectively LLMs support patients with surgical queries. ChatGPT more consistently met expert standards for clarity and thoroughness. As LLM usage expands, ChatGPT may aid patient education on hip surgery, supporting consultations, informed decisions and postoperative guidance.

TP Davis, B. Guevel, K. Logishetty et al. · 0 citations
Open access Jul 2026

Large language models as sources of patient information on robotic knee arthroplasty: a comparative evaluation.

BACKGROUND Robotic-assisted total knee arthroplasty (rTKA) is increasingly used because of its surgical precision. However, inconsistent outcomes and high costs often lead patients to seek additional information from artificial intelligence (AI) tools. Large language models (LLMs) such as ChatGPT-4o, Gemini-2.5-Flash, and DeepSeek-V3 are commonly used, but their reliability and readability in orthopaedics remain unclear. OBJECTIVES To compare the reliability, usefulness, quality, and readability of responses to common patient questions about rTKA generated by leading LLMs. METHODS Three LLMs answered 20 frequently asked patient questions (n = 20) identified through Google Trends and expert validation. Three orthopaedic specialists (n = 3) evaluated reliability, usefulness, and overall quality using validated scales, while readability was assessed with standard indices. RESULTS Inter-rater reliability was good to excellent (ICC = 0.728-0.879). Gemini-2.5-Flash achieved significantly higher reliability and usefulness scores than ChatGPT-4o and DeepSeek-V3 (all p < 0.05). ChatGPT-4o and DeepSeek-V3 produced more readable but less accurate content, revealing an inverse relationship between reliability and readability. CONCLUSIONS Gemini-2.5-Flash provided the most reliable responses, highlighting the need for supervised integration of LLMs in patient education.

Mehmet Utku Çiftçi, A. Koluman, Ebru Aloğlu Çiftçi et al. · 0 citations
Open access Aug 2026

Are large language models such as ChatGPT, capable of supporting patients and general practitioners after spine surgery?

LLMs can support communication and education following spine surgery when used with structured prompting when used with structured prompting and ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers.

S. Wegmann, T. Rosenkranz, Philipp Egenolf et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.