AI-Assisted Assessment Tools: Improving Accuracy and Reducing Bias in Community Medicine and Medical Training Evaluation
Background: Conventional medical exams are usually biased; AI-based applications have the potential to improve grading by decreasing discrimination. The purpose of the study was to compare the traditional examiner-based assessment and AI-aided assessment in community medicine. Methods: This cross-sectional observational study included 120 undergraduate medical students (second and third year) from the Department of Community Medicine, AMS, JSMU, BDMC Sindh, Pakistan between February and June 2024. The respondents were engaged in traditional and AI-assisted tests, based on objective structured clinical examination stations and written exams. MedAssess was used to provide AI scoring. The primary outcomes were mean scores and inter-rater reliability (ICC 0.91); the secondary outcomes were the score differences between genders and the student’s perception on fairness through a validated Likert-scale questionnaire. The analyses were conducted through paired t-tests, ICC analysis, and chi-square tests (p < 0.05). Results: AI-aided evaluation showed higher mean scores (74.6 ± 8.2 vs 71.3 ± 9.5, p = 0.021), better reliability (ICC: 0.91 vs 0.74, p < 0.001) than traditional assessment. The gender differences were also minimized under AI assessment (male: 69.8 ± 9.9 - 74.3 ± 8.4, p = 0.03; female: 72.6 ± 8.8 - 74.8 ± 8.1, p = 0.64). Most of the students believed that AI assessment was fair (82%), trustworthy (76%), and supportive of a hybrid model (88%), (p < 0.001). Conclusion: AI-based evaluation has a positive impact on the consistency of scoring results, bias reduction, and is well-perceived among students. The incorporation of AI platforms into medical teaching can enhance the quality of the assessment.