Predictive Analysis of Learning Behaviour Characteristics for Improving College Students’ Information Literacy using Machine Learning
Abstract
Today, college students must be able to work on both their own and their peers' problems through self-learning and problem-solving skills. However, most current educational tracking programs are unable to accurately identify students who are experiencing difficulty; thus, the number of students that are detected as needing additional academic support is often quite low. The goal of this project is to develop a predictive model for determining the effects of students' information literacy behaviours on students' overall academic success. In order to do this, a dataset of approximately 320 students was collected and the Pearson Correlation method was used to identify the relationships between students' information literacy behaviours and their subsequent learning results. In addition, we used a Comparison of Supervised Learning Algorithms (Decision Tree, Naïve Bayes, K-Nearest Neighbour, Neural Networks, and, Random Forests) to determine which algorithm produced the best results using Python programming. Initially, the Random Forest model produced the best performance with a measurement of 92.50% accuracy and an F1-Score of 89.39%. The results of the study indicate that the use of this data-driven model will help us detect students who are not achieving at an acceptable level.