Skip to content
Open access

Performance of Distance Metrics in SMOTE for Binary Imbalanced Classification

Jul 2026 · Computer Science (CO-SCIENCE) · Vol 6, pp. 134-143 · 0 citations · 30 references

TL;DR

Examination of the choice of distance measure inside SMOTE suggests that the distance function chosen within SMOTE shapes the quality of the generated synthetic points and, in turn, the behavior of the trained classifier.

Abstract

Skewed class distribution continues to be one of the central obstacles in binary classification, since a learning model tends to lean toward the dominant class and consequently overlooks observations belonging to the under-represented class. The purpose of this research is to examine how the choice of distance measure inside SMOTE, specifically Euclidean, Manhattan, Chebyshev, and Hamming, affects predictive quality on imbalanced binary data. Ten publicly available binary datasets drawn from the KEEL repository, whose imbalance ratios span from 1.86 up to 15.80, were used in the experiment. Every dataset was preprocessed and partitioned into 80% for training and 20% for testing; oversampling with SMOTE was carried out on the training portion only, after which four learners, namely Naive Bayes, Decision Tree, Logistic Regression, and k-Nearest Neighbor, were assessed. Model quality was judged through the Matthews Correlation Coefficient (MCC) together with the G-Mean, as these two indicators describe imbalanced performance more faithfully than plain accuracy. The comparison revealed that pairing Euclidean-based SMOTE with Logistic Regression yielded the strongest average scores (MCC = 0.72; G-Mean = 0.79); Manhattan-based SMOTE reached its top MCC again with Logistic Regression (MCC = 0.68) and its top G-Mean with the Decision Tree (G-Mean = 0.79); Chebyshev-based SMOTE delivered the best overall combination together with the Decision Tree (MCC = 0.74; G-Mean = 0.84); and Hamming-based SMOTE performed best alongside Logistic Regression (MCC = 0.73; G-Mean = 0.81). Taken together, these outcomes suggest that the distance function chosen within SMOTE shapes the quality of the generated synthetic points and, in turn, the behavior of the trained classifier.

Read PDF

Similar papers

Open access Sep 2026

Performance Enhancement on Classification of Imbalanced Data using eXtreme Gradient Boosting (XGBoost)

In real world applications, data are generated with uneven distribution called imbalance data which consists of majority and minority classes. For imbalanced data most of the classifier is biased towards the majority class. This means the classifier provides good accuracy for the majority class but very poor accuracy f...

Amit Rauniyar, Prem Chandra Roy, Anisha Pokhrel et al. · 0 citations
Open access Sep 2026

Performance Analysis of Ensemble Technique for Classifying Imbalanced Dataset Using SMOTE-TOMEK Links Sampling

This century has seen a substantial increase in the value of data. Many decisions are made based on data. More the data, more possibility to get accurate result. But in many cases, available datasets are imbalanced i.e. one class have more data and other class have few data which leads to inaccurate classification whil...

Ajaya Shrestha, Dhiraj Pyakurel, Amit Rauniyar · 0 citations
Open access 2026

Optimizing diabetes prediction in machine learning models: Evaluating the effectiveness of a novel class imbalance technique—adaptive synthetic class balancing with class proportion filtering

The skewedness of results when predicting diabetes is mostly due to uneven distribution of data, especially in reducing detection rates of real patients. These are errors which cause delay in treatment or incorrect diagnosis. This work suggests a counter plan to this assumption, which is Adaptive Synthetic Class Balanc...

Pankaj Beldar, Snehal M. Kamalapur, Priti Vaidya et al. · 0 citations
Conference Open access 2026

Supervised Learning for Classification in Data Science: A Comparative Perspective

Overall, logistic regression demonstrates strong and good generalization ability, while ensemble methods constitute robust alternatives for binary classification on structured data.

Zahra Benider, H. Bouzahir, Jaafar Idrais · 0 citations
Review Sep 2026

EVALUATION OF TEXT CLASSIFICATION ALGORITHMS: A STUDY ON CLASS IMBALANCE AND PREPROCESSING

This paper presents a controlled comparative study of traditional machine learning algorithms for thematic text classification, focusing on the impact of preprocessing strategies and class imbalance on model performance under different data conditions. Two experimental scenarios were considered: the Women’s Clothing E-...

Dulce Liliana Estrada Bahena, Alicia Martínez Rebollar, Hugo Estrada Esquivel et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.