Examination of the choice of distance measure inside SMOTE suggests that the distance function chosen within SMOTE shapes the quality of the generated synthetic points and, in turn, the behavior of the trained classifier.
Abstract
Skewed class distribution continues to be one of the central obstacles in binary classification, since a learning model tends to lean toward the dominant class and consequently overlooks observations belonging to the under-represented class. The purpose of this research is to examine how the choice of distance measure inside SMOTE, specifically Euclidean, Manhattan, Chebyshev, and Hamming, affects predictive quality on imbalanced binary data. Ten publicly available binary datasets drawn from the KEEL repository, whose imbalance ratios span from 1.86 up to 15.80, were used in the experiment. Every dataset was preprocessed and partitioned into 80% for training and 20% for testing; oversampling with SMOTE was carried out on the training portion only, after which four learners, namely Naive Bayes, Decision Tree, Logistic Regression, and k-Nearest Neighbor, were assessed. Model quality was judged through the Matthews Correlation Coefficient (MCC) together with the G-Mean, as these two indicators describe imbalanced performance more faithfully than plain accuracy. The comparison revealed that pairing Euclidean-based SMOTE with Logistic Regression yielded the strongest average scores (MCC = 0.72; G-Mean = 0.79); Manhattan-based SMOTE reached its top MCC again with Logistic Regression (MCC = 0.68) and its top G-Mean with the Decision Tree (G-Mean = 0.79); Chebyshev-based SMOTE delivered the best overall combination together with the Decision Tree (MCC = 0.74; G-Mean = 0.84); and Hamming-based SMOTE performed best alongside Logistic Regression (MCC = 0.73; G-Mean = 0.81). Taken together, these outcomes suggest that the distance function chosen within SMOTE shapes the quality of the generated synthetic points and, in turn, the behavior of the trained classifier.
In real world applications, data are generated with uneven distribution called imbalance data which consists of majority and minority classes. For imbalanced data most of the classifier is biased towards the majority class. This means the classifier provides good accuracy for the majority class but very poor accuracy f...
Amit Rauniyar, Prem Chandra Roy, Anisha Pokhrel et al.· Journal of Hillside College...· 0 citations
This century has seen a substantial increase in the value of data. Many decisions are made based on data. More the data, more possibility to get accurate result. But in many cases, available datasets are imbalanced i.e. one class have more data and other class have few data which leads to inaccurate classification whil...
Ajaya Shrestha, Dhiraj Pyakurel, Amit Rauniyar· Journal of Hillside College...· 0 citations
The pipeline approach ensured no data leakage in cross-validation, and the findings support ensemble ML models with SMOTE as a preprocessing step for imbalanced CVD datasets.
M. Maindarkar, M. Patil, Pratibha Jadhav et al.· Journal of Intelligent Decis...· 0 citations
The skewedness of results when predicting diabetes is mostly due to uneven distribution of data, especially in reducing detection rates of real patients. These are errors which cause delay in treatment or incorrect diagnosis. This work suggests a counter plan to this assumption, which is Adaptive Synthetic Class Balanc...
Pankaj Beldar, Snehal M. Kamalapur, Priti Vaidya et al.· Sigma Journal of Engineering...· 0 citations
Overall, logistic regression demonstrates strong and good generalization ability, while ensemble methods constitute robust alternatives for binary classification on structured data.
Zahra Benider, H. Bouzahir, Jaafar Idrais· EPJ Web of Conferences· 0 citations
This paper presents a controlled comparative study of traditional machine learning algorithms for thematic text classification, focusing on the impact of preprocessing strategies and class imbalance on model performance under different data conditions. Two experimental scenarios were considered: the Women’s Clothing E-...
Dulce Liliana Estrada Bahena, Alicia Martínez Rebollar, Hugo Estrada Esquivel et al.· DYNA· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.