Large Language Models for Ankle Fracture Classification and Management Prediction from Routine Clinical Documentation: A Single-Center Exploratory Study.
Sep 2026· Journal of imaging informatics in medicine· 0 citations· 23 references
Medicine
TL;DR
It is suggested that LLMs can extract structured information from routine clinical documentation, performing well for standardized classification but less reliably for procedure-level prediction.
Abstract
Ankle fractures are among the most common injuries in trauma surgery and require accurate classification, consistent documentation, and individualized management. Large language models (LLMs) offer the potential to translate unstructured clinical text reports into structured, management-related information, yet their role in orthopedic trauma workflows remains insufficiently defined. In this retrospective study, the performance of Llama 3.1 (70B Instruct) was evaluated using routine radiology reports and clinical documentation from 54 patients with acute ankle fractures. Outputs were compared with reference standards derived from routine clinical documentation for Weber fracture classification, operative versus nonoperative management, and surgical procedure category. Four prompting strategies were systematically assessed. Accuracy for Weber classification ranged from 0.759 to 0.815, well above majority-class baseline (0.556), with strong performance for Weber B fractures and most errors occurring between adjacent categories. Macro-averaged F1 across the three Weber classes ranged from 0.744 to 0.822. Operative management prediction reached accuracies of 0.833 to 0.963 (sensitivity 0.872 to 1.000, specificity 0.571 to 0.714). Procedure category prediction demonstrated lower accuracy (0.426 to 0.610 depending on the endpoint), largely at or below a majority-class baseline (0.553), reflecting the limitations of text-only input for operative planning. These findings suggest that LLMs can extract structured information from routine clinical documentation, performing well for standardized classification but less reliably for procedure-level prediction. Prospective and multimodal validation is needed before clinical implementation.
Pediatric physeal fractures can be difficult to assess radiographically. We characterized regional fracture-detection failure patterns of a multimodal large language model (MLLM) in confirmed pediatric physeal fractures.
This retrospective, single-center, case-only study included 805 confirmed physeal frac...
S. E. Erginoğlu, N. Ülgen, Ali Said Nazlıgül et al.· BMC Medical Imaging· 0 citations
Multimodal large language models can interpret medical images, but their performance for pediatric elbow radiographs remains uncertain. We evaluated the diagnostic performance of GPT-5.2 Instant as accessed through the ChatGPT web interface during the defined study period.
In this prospective, single-cente...
O. Taş, Mehmet Yorgun, R. Aktaş et al.· BMC Medical Imaging· 0 citations
Objective
To evaluate the diagnostic performance of multiple large language models (LLMs) against expert consensus in determining surgical intervention needs for feline metacarpal and metatarsal fractures.
Methods
In this retrospective study (December 2023 to February 2025), 73 clinical cases of feline metacarpal and...
S. Okur, Ç. Özkalıpçı, Büşra Baykal et al.· Journal of the American Vete...· 0 citations
Frontal sinus fracture management requires integration of aesthetic, sinonasal, and intracranial risks, making it a demanding test of artificial intelligence-generated clinical advice. This study compared ChatGPT, Gemini, and DeepSeek across 30 systematically developed frontal sinus fracture vignettes. Ninety model-vig...
Mehmet Sefik Oruc, A. Gunenc, Ovunc Akdemir et al.· The Journal of craniofacial...· 0 citations
It is found that structured prompt engineering improves off-the-shelf LLM performance over unstructured prompting across age groups and enables general-domain LLMs to achieve clinician-comparable performance in pediatric trauma triage across age groups.
Brendan Fox, A. Balaji, Philip J. Seger et al.· Journal of Pediatric Surgery· 0 citations
Femoral neck fractures are a heterogeneous entity with wide variability in patient profiles and outcomes. Single-variable classification may not capture the multidimensional interactions that influence prognosis. This study aimed to identify clinically meaningful patient subgroups using hierarchical cluster analysis an...
Enver Ipek, Yusuf Altuntaş, Bahadır Balkanlı et al.· Archives of Orthopaedic and...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.