LARGE LANGUAGE MODELS AS CLINICAL TUTORS: AN EXPLORATORY PILOT STUDY OF FEASIBILITY AND ACCEPTABILITY IN AUTOMATED CLINICAL NOTE EVALUATION AND STUDENT FEEDBACK
Abstract
Introduction and Objective: Manual evaluation of medical histories in medical education faces challenges such as limited preceptorship time, insufficient feedback, and subjectivity in assessment. This exploratory pilot study evaluated the feasibility, acceptability, and preliminary impact of using a Large Language Model (LLM) as a clinical tutor for automated medical history evaluation and feedback to medical students. Methodology: This observational, longitudinal, retrospective study used quantitative and qualitative approaches. Five fourth-semester medical students participated in seven evaluation cycles, totaling 35 medical histories. The intervention used the NotebookLM platform structured with the Chain-of-Thought technique. The LLM assessed clinical rigor, applied an evidence-based checklist, and provided feedback using positive reinforcement. All reports were reviewed by a senior preceptor before being sent to the students. A focus group was conducted at the end of the study. Results: Scores improved from 6.0-8.0 (mean±SD: 7.0±0.82) to 9.6-10.0 (9.88±0.13). Participants reported increased confidence and good acceptance of the structured feedback. The preceptor modified 14 of the 35 reports (40%), mainly to correct algorithmic rigidity and add clinical context. Conclusion: This study demonstrates the technical feasibility and preliminary acceptability of LLM-based clinical tutoring when combined with faculty supervision (Human-in-the-Loop).