Aug 2026· Bariatric Surgical Practice and Patient Care· 0 citations· 11 references
TL;DR
Strategic model selection can enhance preoperative workflows by balancing comprehensiveness with targeted, actionable assessment, however, deployment should remain clinician-supervised because outputs may vary with model updates, prompt phrasing, and local workflow constraints.
Abstract
Psychological assessment is essential for bariatric surgery candidacy but remains inconsistent and time-consuming. This study compared three large language models (LLMs) in generating structured psychological screening checklists to support standardized preoperative evaluation.
Models A, B, and C generated checklists for five standardized bariatric vignettes. Three blinded psychologists independently rated the outputs on 5-point Likert scales for clinical relevance, completeness, specificity, and usability. Statistical analysis included Jaccard similarity to quantify content overlap and analysis of variance with Tukey
post hoc
tests to compare expert ratings.
Model A produced the most extensive checklists (mean = 35.6 items, standard deviation = 4.2), while Models B and C were more concise (28.4 and 26.7 items, respectively). Model A achieved the highest completeness scores, whereas Model B was rated highest for specificity and usability. Interrater agreement was good to excellent (intraclass correlation coefficient = 0.79–0.87). Moderate content overlap (Jaccard similarity = 0.45–0.63) suggested complementary model strengths. Between-model differences were significant for completeness,
F
(2, 6) = 11.5,
p
= 0.008, and specificity,
F
(2, 6) = 9.8,
p
= 0.013.
LLMs differ significantly in checklist breadth and focus. Strategic model selection can enhance preoperative workflows by balancing comprehensiveness with targeted, actionable assessment. However, deployment should remain clinician-supervised because outputs may vary with model updates, prompt phrasing, and local workflow constraints.
Introduction: Surgical oral board examinations are vulnerable to examiner variability, potentially compromising scoring reliability. This study evaluated six public frontier large language models (LLMs) as both examinees and graders on general surgery oral board-style cases. Methods: Six commercially available LLMs were assessed across three standardized cases using a fully crossed psychometric design. Responses were graded on a 0-3 ordinal scale by six LLMs and three blinded senior surgeons using a structured rubric with anchored descriptors for each score level. Analyses included performance comparisons, reliability testing, and variance decomposition. Results: All LLMs achieved passing or near-passing performance. Significant differences existed between models (Friedman χ² = 12.51, p = 0.028). AI graders demonstrated greater internal consistency than the three-surgeon human panel in this pilot study (Cronbach's α = 0.697 vs. −0.923), though this comparison is limited by the small human panel and the absence of rater calibration for human graders. The dominant source of score variance was the Examinee × Rater interaction, suggesting that grader-driven disagreement accounted for more score variability than actual examinee performance differences in this dataset. Conclusions: LLMs demonstrate passing-level oral board performance and more consistent grading than human surgeons in this simulation context, suggesting potential value for further investigation of LLM integration into surgical training and assessment frameworks, pending replication in larger samples. Future work should evaluate whether these findings generalize to live examination settings, compare against calibrated human rater panels, and explore AI-augmented panel designs.
Kian A. Huang, Haris K. Choudhary, A-Lan Xu et al.· Cureus· 0 citations
This study aimed to compare the accuracy, completeness, and readability of responses generated by three large language models (LLMs)—ChatGPT-4.0 (OpenAI), Microsoft CoPilot, and Google Gemini—regarding the treatment and management of idiopathic congenital talipes equinovarus (ICTEV) using the Ponseti method.
Fifteen frequently asked questions were selected from pediatric orthopedic center websites, Google Trends analysis, and clinical experience. Each question was submitted verbatim in a new session to the three LLMs within 24 h. Seven board-certified pediatric orthopedic surgeons, blinded to the source, rated responses for accuracy (5-point Likert scale) and completeness (3-point Likert scale). Readability was assessed using the Flesch–Kincaid grade level. Mean scores ± standard deviation were calculated, and inter-rater reliability was estimated using the intraclass correlation coefficient (ICC). Group differences were tested with ANOVA and chi-squared tests (
p
< 0.05).
A total of 45 responses were evaluated. Gemini achieved the highest mean accuracy (4.1 ± 0.8), followed by CoPilot (3.6 ± 0.8) and ChatGPT-4.0 (3.3 ± 0.9), with significant differences among models (
p
< 0.001). Completeness ratings also differed significantly (
p
< 0.001). Readability analysis showed that ChatGPT produced shorter, more readable text, while Gemini generated longer, more complex responses. Inter-rater reliability was substantial for accuracy (ICC 0.715) and completeness (ICC 0.710).
Google Gemini outperformed ChatGPT-4.0 and CoPilot in accuracy and comprehensiveness for ICTEV management using the Ponseti method. However, its complex responses may limit accessibility, whereas ChatGPT-4.0 offered more readable but less detailed answers.
IV
A. Vescio, G. Testa, M. Sapienza et al.· Journal of Children's Orthop...· 0 citations
Purpose This study was aimed to compare the efficacy of three most popular large language models (LLMs)—Claude Opus 4.6, ChatGPT Thinking 5.4 and DeepSeek v3.2 in answering frequently asked questions (FAQs) about scoliosis. Methods 20 scoliosis related questions (four categories, five questions in each category) were submitted to each LLM. A panel of 9 experts (two spine surgeons, two pediatric orthopedic surgeons and five physical therapists, all blinded to the LLMs and responses) rated independently each response generated by LLMs on a 6 points Likert scale (1 as strongly disagree to 6 as strongly agree). 540 total ratings were collected. Intergroup comparisons were conducted by Kruskal Wallis test and Mann Whitney U pairwise tests. Paired question level analysis was achieved by Friedman test and Wilcoxon signed rank comparisons. Results Claude’s score was 5.53 ± 0.76 much higher than both ChatGPT (4.84 ± 0.86, p < 0.001) and DeepSeek (4.86 ± 0.84, p < 0.001), but no difference was found between ChatGPT and DeepSeek (p = 0.749). Claude performed on top for 19 of 20 questions (95%) and was favored by 7 of 9 reviewers. Consistency of Claude was also highest [CV = 13.8% vs. 17.9% (ChatGPT) and 17.2% (DeepSeek)]. Conclusion Although all three LLMs achieved favorable overall ratings (>4.8/6), Claude performed significantly better than ChatGPT and DeepSeek for scoliosis FAQs taking into account its higher accuracy and consistency. Within the scope of the present evaluation, Claude demonstrated the strongest overall performance among the three LLMs tested.
Yi-Chen Wang, Xuanyan Liu, Mengjie Chen et al.· Frontiers in Public Health· 0 citations
Background: Hypertension is a chronic condition significantly impacting physical and psychological well-being. HRQoL instruments must undergo cross-cultural adaptation and psychometric evaluation before clinical application. Objective: This systematic review evaluated the psychometric robustness (reliability, validity, responsiveness) of hypertension-specific HRQoL instruments. Methods: Eight databases were searched for original research published in English or Indonesian (2021–2026) following PRISMA 2020. Methodological quality and evidence certainty were assessed using COSMIN and modified GRADE. Results: 13 studies from 10 countries were included. Six instruments (46.1%) demonstrated strong psychometric profiles, including adequate content validity and robust reliability. While content validity (100%) and internal consistency (84.6%) were strong due to rigorous adaptation, divergent validity and responsiveness were severely underreported (84.6%). According to COSMIN guidelines, test-retest reliability studies with sample sizes below 50 participants are considered to have 'Low' GRADE evidence quality due to insufficient precision, despite demonstrating adequate ICC values. QLICD-HY, STI-HRQoL, and DASI provided the strongest evidence. Conclusion: Instruments are generally reliable, yet future research must evaluate clinical responsiveness through longitudinal designs. Instruments with the strongest psychometric profiles, such as QLICD-HY and DASI, are recommended for accurate patient assessment.
Afif Rakha Murtadha, T. Andayani, Dwi Endarti· Journal of Pharmacy and Scie...· 0 citations
The Kihon Checklist (KCL) is a brief, multidimensional, and function-oriented instrument for frailty screening. However, a comprehensive synthesis of its psychometric properties remains lacking. This scoping review aimed to map and summarize psychometric evidence on the reliability and validity of the KCL in frailty screening. A systematic search of nine databases identified 22 eligible English-language studies published up to May 31, 2025. Most studies were conducted in community settings and findings on internal consistency reliability and concurrent validity were generally favourable. Evidence for predictive validity was mixed, while content validity was primarily evaluated in cross-cultural adaptation studies. Limited evidence was found for test-retest reliability, split-half reliability, and structural, convergent, or known-groups validity. Overall, the findings suggest that the KCL shows promise as a psychometrically evaluated instrument for frailty screening, especially in community settings. Further research is needed to address current evidence gaps and expand its applicability across diverse populations and clinical contexts.
Yang Zhao, Ting-Ting Wang, Ling-Na Kong et al.· Geriatric Nursing· 0 citations
BACKGROUND
Standard setting in Objective Structured Clinical Examinations (OSCEs) typically identifies a single cut score to distinguish competent from non-competent candidates. However, this approach does not help defensible differentiation between 'competent' and 'excellent' students.
OBJECTIVE
This study introduces a Two-Point Regression model, an extension of the borderline regression method, to establish two cut scores, one for competence and one for excellence, enabling more nuanced classification of student performance.
METHODS
A retrospective analysis was performed on 10 OSCEs delivered to UK medical students. For each station, linear regressions of global ratings (0-3) against checklist scores were used to generate the traditional pass cut score (cX = 1). Additional thresholds for excellence were generated at global thresholds eX = 2, 2.25, 2.5 and 3. Outcomes included: proportion graded as excellent (≥70% after scaling), mean scaled grades, and mean global scores of 'excellent' students. Statistical analyses included chi-square, ANOVA, and post-hoc testing as appropriate.
RESULTS
Pass/fail classifications were unchanged across all methods. The Two-Point regression model substantially reduced the proportion of students achieving excellence from 47% to 17.2% when the excellence cut score was set to a global threshold of eX = 2.25, aligning with national benchmarks. Mean global scores of 'excellent' students increased in line with threshold choice, demonstrating strong internal consistency (R2=0.9939).
CONCLUSION
The Two-Point Regression method offers an easily implementable and defensible approach for identifying excellence in numerically graded OSCEs, improving alignment between examiner global judgements and awarded grades. Threshold choice should be context-dependent, but the Two-Point Regression method provides a robust and adaptable framework.
A. Lunn, Christopher J. Harrison, J. McLachlan· Medical Teacher· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.