Jul 2026· Int. J. Medical Informatics· Vol 220, pp.
106589
· 0 citations· 21 references
MedicineComputer Science
TL;DR
While LLM use among surveyed medical students in Canada was nearly universal and perceived favorably, students reported exposure to inaccurate outputs and substantial gaps in formal training and privacy literacy.
Abstract
Background
Large language models (LLMs) are increasingly embedded in medical education and clinical care settings, yet contemporary Canadian data describing medical students' use and perceptions remain limited.
Objective
To quantify the prevalence, frequency, and patterns of LLM use among medical students in Canada; to characterize perceptions of utility, accuracy, limitations, and impact; and to describe perceived barriers, challenges, and ethical/privacy concerns.
Methods
We conducted a national, cross-sectional survey distributed to English-speaking medical students between November and December 2025. Recruitment occurred through medical school channels, student unions, and national/regional student organizations.
Results
Among 286 respondents from 10 medical schools, 96.50% reported using at least one LLM. The most commonly used LLMs were ChatGPT (93.36%) and OpenEvidence (57.69%). Daily/weekly use was most frequent for coursework assistance (60.22%) and clinical questions (57.14%). Most respondents reported positive impacts on efficiency (81.62%), learning (77.01%), and academic performance (59.49%). Students commonly reported encountering inaccurate information (90.18%). Formal instruction on LLM use was uncommon (10.95%), though 67.67% of students agreed medical schools should integrate formal instruction on LLMs. Only 21.43% of respondents felt adequately educated on data privacy regulations applicable to these tools.
Conclusion
While LLM use among surveyed medical students in Canada was nearly universal and perceived favorably, students reported exposure to inaccurate outputs and substantial gaps in formal training and privacy literacy. These findings support the development of structured curricular guidance on appropriate application of these tools, including information verification practices and ethical, privacy-aware engagement.
BACKGROUND
Large language model (LLM)-generated hospital courses are increasingly integrated into electronic health records (EHRs), yet their accuracy and safety in pediatric populations remain poorly characterized.
OBJECTIVE
To evaluate the accuracy, text quality, and perceived potential harm of EHR-integrated and LLM-generated hospital courses in pediatric inpatient care during early clinical implementation.
METHODS
We conducted a descriptive evaluation from June 10 to August 8, 2025, at an academic freestanding children's hospital using an Epic EHR with an integrated LLM tool (GPT-4o and GPT-4.1). Clinicians across multiple roles, including attending physicians, residents, and advanced practice providers, reviewed LLM-generated hospital courses for their own patients. Clinicians identified and categorized errors (hallucinations, inaccuracies, or omissions). They also rated text quality (comprehensiveness, conciseness, coherence) on a 5-point scale and perceived harm on an 8-point scale.
RESULTS
A total of 129 LLM-generated hospital courses were reviewed (median length of stay, 3 days; IQR, 2-7) by 50 involved clinicians. Hallucinations occurred in 21% (95% CI, 14%-29%) of the hospital courses, inaccuracies in 41% (53/129; 95% CI, 33%-50%), and omissions in 24% (31/129; 95% CI, 17%-32%). Overall, perceived harm ratings were low (median, 0; IQR, 0-1). Text quality ratings were high (median [IQR]: comprehensiveness, 4 [3-5]; conciseness, 4 [4-5]; coherence, 4 [4-5]) and comparable with prior literature.
CONCLUSION
In this pediatric evaluation of LLM-generated hospital courses reviewed by frontline clinicians, errors were common, but perceived potential harm was low, even assuming use without clinician correction. These findings support the use of LLM-generated hospital courses as starting drafts when paired with clinician review and institutional safeguards.
Jasmine E. Kim, Jonathan D. Hron, Daniel J. Kats et al.· Hospital Pediatrics· 0 citations
ObjectiveThis study evaluates the performance of large language models (LLMs)-ChatGPT-4.0, Gemini 2.0 Pro, o3-mini, Doctor GPT and DeepSeek-V3-in a national orthopaedic proficiency examination and explores their implications for health informatics and medical education. The responses of these models were analysed to assess accuracy rates and differences between models.MethodA total of 100 multiple-choice questions from the 2024 TOTEK examination were administered to each AI model under identical conditions. Correct and incorrect responses were recorded, and differences in performance were evaluated using chi-square testing and frequency analysis. Question categories were also compared to identify domain-specific variations.Resultso3-mini achieved the highest accuracy rate (79%), while Gemini 2.0 showed the lowest (68%); all models exceeded the 60% pass threshold. A statistically significant difference between models was identified in the Surgical Procedures category, in which Gemini 2.0 answered fewer questions correctly (23/36) than the other models (30-32/36) (χ2 = 9.87, df = 4, p = 0.043). No significant differences were observed in the remaining categories (all p > 0.05), and the overall difference in accuracy between models did not reach statistical significance (χ2 = 4.01, df = 4, p = 0.405). Clinical decision-making and visual content-based questions were the most challenging for all models.ConclusionAI models demonstrate generally high accuracy in medical examinations; however, they struggle with interpreting clinical context, recognising atypical medical scenarios and answering questions involving visual content.
Bünyamin Arı· Health Informatics Journal· 0 citations
Background: Artificial intelligence (AI) is increasingly being used in medical education, offering opportunities for personalized learning, rapid access to information, and support for academic activities. However, concerns regarding accuracy, overdependence, and reduced critical thinking remain.
Objectives: To assess the perceptions, attitudes, perceived benefits and concerns regarding the use of AI tools among undergraduate medical students.
Methods: A cross-sectional study was conducted among 767 undergraduate medical students at a tertiary care teaching institution in Bharuch, Gujarat, India, from January to March 2026. Students from all MBBS years who provided informed consent were included. Data were collected using a pre-tested, semi-structured questionnaire administered through Google Forms. Descriptive statistics were used to summarize the findings as frequencies and percentages.
Results: Among 767 participants, AI use was reported daily by 207 (27.0%) students, on a few days per week by 282 (36.8%), and occasionally or rarely by 278 (36.2%). ChatGPT was the most commonly used AI tool reported by 698 (91.0%) students, followed by Google Gemini 398 (51.9%). The most common academic purpose of AI use was clarifying difficult concepts, reported by 561 (73.1%) students, followed by obtaining quick summaries or notes 460 (60.0%). AI was considered very helpful for understanding concepts by 445 (58.0%) students, while 515 (67.1%) reported faster learning. Nearly half 367 (47.8%) supported integrating AI into medical teaching. However, inaccurate information 334 (43.4%), reduced critical thinking 327 (42.6%), and overdependence 252 (32.8%) were major concerns.
Conclusions: AI tools are widely used and positively perceived by medical students. Structured and responsible integration of AI into medical education may maximize its benefits while minimizing potential risks.
Keywords: Artificial Intelligence, ChatGPT, Medical Education, Medical Students
Vaishali Patel, Vallari Jadav, Kuntal Patel· International journal of sci...· 0 citations
Although uptake was high, trust was moderate, and students expressed concerns about professionalism and critical thinking, medical schools should provide explicit guidance on acceptable use, incorporate artificial intelligence literacy training and ethical use guidelines, and redesign assessment to protect the skills students perceive to be most at risk.
Eunah Joo, Oliver B Ma, T. Woolley et al.· International Medical Educat...· 0 citations
The creation of new medical schools has presented a challenge to providing high quality medical education in UK hospitals. A second medical school cohort was introduced to a large London hospital in September 2024, with a further year group joining in September 2025. This presented an opportunity to evaluate the impact on student learning.
Evaluation of medical student perceptions of learning was undertaken through a single centre cross-sectional study. A 15-component questionnaire was distributed to all medical students on placement over a two-week period in November 2025. The questionnaire evaluated all barriers to learning, rather than explicitly mentioning the introduction of a new medical school, to avoid bias.
35 students responded to the survey with 80% responses from one medical school and 20% from the other (approximately reflecting overall cohort numbers). The results showed high student satisfaction with 74% of students stating they ‘strongly agreed’ with the phrase “I had enough teaching to meet my learning needs”. Where students cited barriers to learning, only 2.5% of students reported this being due to the presence of medical students from a different university.
This study presents evidence that medical student perception of the quality of their clinical experience and learning was generally positive, despite the recent introduction of a second school. This situation is likely to become more common given the increasing number of medical school places; the findings of this study may help alleviate fears that this will have an inherently negative impact on student learning.
Aslesha Deshraju, Olivia Flood, K. Small et al.· British Journal of Surgery· 0 citations
BACKGROUND
With the growing integration of generative Artificial Intelligence (AI) into healthcare, the DeepSeek large language model has emerged as a versatile tool for medical students, offering adaptive learning solutions tailored to various medical scenarios. However, the adoption of AI in medicine also raises ethical concerns that require warrant consideration. This study aimed to investigate medical students' usage patterns, attitudes, and influencing factors related to DeepSeek.
METHODS
In July 2025, a cross-sectional survey was conducted using random sampling among 1000 medical students from institutions affiliated with Anhui Medical University. Based on previous AI attitude scales, a validated self-administered questionnaire was used to collect data on students' attitudes, DeepSeek usage patterns.
RESULTS
Among the 937 valid responses (response rate: 93.7%), 874 participants were aware of DeepSeek prior to the survey and 765 had used it. The Cronbach's alpha value for the questionnaire was 0.83. 42.0% of students doubted the accuracy of information provided by DeepSeek, 48.6% were apprehensive about potential plagiarism accusations, and 43.0% worried about over-reliance on AI. Additionally, 46.6% found DeepSeek interesting and appealing, with 37.7% expressing enthusiasm for learning new AI technologies. Several factors were significantly associated with usage experience, including being female [β = 6.1(1.5-10.7), (p = 0.031)], holding a Master's degree [β = 10.5(2.1-18.1), (p = 0.044)], and engaging in clinical [β = 7.4(1.2-13.6), (p = 0.021)] or basic research [β = 9.2(2.9-15.4), (p = 0.012)]. Regarding attitudes, significant predictors included a Master's degree [β = -2.9(-6.2 to-0.5), (p = 0.011)], having a stomatology background [β = 8.0(2.3-13.7), (p = 0.035)], engaging in clinical [β = -2.6(-5.0 to -0.1), (p = 0.036)] or basic [β = -2.7(-5.2 to -0.3), (p = 0.042)] research.
CONCLUSIONS
Among the medical students surveyed in this study, 94.8% reported that they would be willing to use and recommend DeepSeek, but they maintain a cautious attitude toward its privacy, reliability, and dependency. Despite these concerns, 90.2% of participants said they would be willing to use DeepSeek to help them complete their tasks.
Wei Liu, Leilei Cao, Xiaoyan Chen et al.· BMC Medical Education· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.