EVALUATION OF TEXT CLASSIFICATION ALGORITHMS: A STUDY ON CLASS IMBALANCE AND PREPROCESSING
Abstract
This paper presents a controlled comparative study of traditional machine learning algorithms for thematic text classification, focusing on the impact of preprocessing strategies and class imbalance on model performance under different data conditions. Two experimental scenarios were considered: the Women’s Clothing E-Commerce Reviews dataset, characterized by a highly imbalanced class distribution, and the AG News dataset, used as a balanced reference dataset. The preprocessing pipeline included text cleaning, stopword removal, tokenization, lemmatization, TF-IDF vectorization, class balancing, and dimensionality reduction. Six models were evaluated: Naive Bayes, Random Forest, Support Vector Machine, Logistic Regression, K-Nearest Neighbors, and K-Means as an unsupervised baseline. The results show that Random Forest achieved the best performance on the imbalanced Clothing dataset, reaching an accuracy of 0.763, an F1-score of 0.777, and an MCC of 0.739. In contrast, K-Nearest Neighbors obtained the highest scores on the balanced AG News dataset, with an accuracy and F1-score of 0.927. Naive Bayes showed limitations when dealing with sparse and ambiguous textual representations. The main contribution of this work is a systematic experimental analysis that highlights how preprocessing decisions, feature sparsity, and class distribution influence the behavior of traditional machine learning models in text classification tasks. These findings provide useful insights for model selection and evaluation under different data scenarios. Key Words: Text Classification; Text Preprocessing; Machine Learning; Text Analysis; Class Imbalance