Evaluation of an Artificial Intelligence–Driven Language Model for Age-Specific Patient Education in Pediatric Orthopaedics
TL;DR
GenAI LLMs can generate age-specific explanations for common surgical procedures, treatments, and diagnostic markers in pediatric orthopaedics using an artificial intelligence language model based on the child's age.
Abstract
Background Describing complex orthopaedic surgical procedures and diagnoses to pediatric patients across a large age range and at varying levels of health literacy can be quite challenging. Generative artificial intelligence (GenAI) offers a potential solution by providing real-time age specific explanations for pediatric patients. The use of GenAI large language models (LLMs) could potentially allow providers to more effectively communicate complex medical problems in an age-appropriate and comprehensible manner. This study aims to validate the readability, accuracy, and comprehensiveness of responses produced by the ChatGPT-4 LLM (OpenAI, San Francisco, CA) when prompted to explain a specific pediatric orthopaedic procedure indicated for a given diagnosis to patients of 3 different age groups. Methods ChatGPT-4 explanations were analyzed for 43 pediatric orthopaedic procedures and tests across prompts for 3 different age groups: 6 years old, 10 years old, and 14 years old. Responses were assessed for readability using 6 quantitative readability indices. The accuracy and comprehensiveness of the responses were rated by 7 board-certified pediatric orthopaedic surgeons, who also compared them to their own explanations used in clinical practice. Results Each of the 6 indices showed a significant difference in readability between the 6-, 10-, and 14-year-old age group prompts (P < .001). The 6-year-old prompts were more readable than the 10- and 14-year-old prompts (P < .001), and the 10-year-old prompts were more readable than the 14-year-old prompts (P < .001) . Out of a 5-point scale, the mean (standard deviation) attending-rated accuracy for the artificial intelligence responses was 4.15 (0.81), with no significant differences in rated accuracy between procedures (P = .051) or age groups (P = .27). The mean attending-rated comfort was 3.95, with inter-rater reliability coefficients for accuracy and comfort of 0.92 and 0.94, respectively. Conclusions GenAI LLMs can generate age-specific explanations for common surgical procedures, treatments, and diagnostic markers in pediatric orthopaedics. Physicians validated the accuracy of its responses and demonstrated a high level of comfort using the GenAI explanations in clinical settings. Key Concepts (1) An artificial intelligence (AI) language model adjusted its explanations based on the child’s age, providing simpler language for younger children and more advanced explanations for older children.(2) Across 6 established readability measures, explanations written for younger children were consistently easier to understand than those created for older age groups.(3) Seven board-certified pediatric orthopaedic surgeons found the AI-generated explanations to be highly accurate, with an average rating of 4.15 out of 5.(4) Physicians were generally comfortable with the use of AI-generated educational explanations in practice, and their ratings were highly consistent with one another.(5) When reviewed by a physician, generative AI could be a helpful tool for improving patient education, health literacy, and shared decision-making in pediatric orthopaedic care. Level of Evidence III/IV