Aug 2026· International Journal of Data Science and Analysis· Vol 22· 0 citations· 35 references
TL;DR
This work proposes a novel hybrid approach that combines the strengths of the extreme learning machine (ELM) and the twin support vector machine (TSVM) to address the challenges of robustness and scalability in multi-label classification, particularly in settings where deep learning is not practical due to limited training instances.
This review consolidates the landscape of CP adaptations for MLL under a unified framework, examining the types of outputs and guarantees they provide, where label dependencies are incorporated, and how inference cost scales with the number of labels.
Multiclass classification is a fundamental problem across a wide range of domains. It is still challenging due to possession of high inter-class similarity, class imbalance datasets, and variability in data distributions. Rule-based classifiers such as XGBoost often achieve stronger performance on structured features, but they are limited in capturing smooth functional relationships among variables. Similarly, neural network models can represent complex nonlinear interactions but frequently suffer from overfitting and generalization issues. To address these limitations, we propose LFS-FRAME, a Leakage-Free Stacked ensemble framework that integrates functional learning using Kolmogorov-Arnold Networks (KAN) and rule-based learning via XGBoost for robust multiclass classification. The proposed framework constructs unbiased meta-features by employing a strict out-of-fold stacking strategy to ensure complete isolation between training and validation data hence preventing performance leakage. By learning over probabilistic outputs from heterogeneous base learners, the meta-classifier effectively exploits both global functional patterns and sharp decision boundaries present in the complex data. Experimental evaluations on multi-class datasets demonstrate that LFS-FRAME improves performance metrics, and overall accuracy is 89.85% in identifying major families and 81.74% in identifying sub-families relative to strong single-model baselines. These results highlight the effectiveness of leakage-free functional and rule-based stacking for reliable and generalizable multiclass classification.
S. P. Sharmila, Aruna Tiwari· arXiv.org· 0 citations
The small-sample-size (SSS) problem remains a fundamental challenge in machine learning when labeled data are scarce due to cost, accessibility, or ethical constraints. While numerous approaches have been proposed, existing methods often struggle to maintain stable and discriminative representations under high-dimensional and limited-data conditions. Kernelized Linear Principal Component Discriminant Analysis (KLPCDA), a recently proposed modular framework, integrates variance preservation, inter-class separability, and intra-class compactness within a unified kernel space. Although its formulation has shown promising initial results, a systematic understanding of how its components interact across diverse SSS scenarios remains lacking. In this paper, we present a systematic cross-domain study of KLPCDA to characterize the interaction mechanisms among its core objectives. We analyze the behavior of its seven variants across multiple real-world SSS tasks, including hyperspectral image classification, mechanical fault diagnosis, medical diagnosis, and face recognition. Through extensive experiments and ablation studies, we investigate how different objective combinations influence performance under varying conditions such as noise, class imbalance, and high dimensionality. Our analysis reveals consistent patterns in the interaction of the three core objectives variance, between-class, and within-class terms, providing a unified and interpretable understanding of their roles in stabilizing representations and enhancing discrimination in SSS settings. Based on these findings, we further derive practical guidelines for selecting appropriate KLPCDA variants under different data characteristics. Experimental results demonstrate that KLPCDA achieves strong and robust performance across domains, while maintaining low computational complexity suitable for resource-constrained environments.
Multi-label classification (MLC) seeks to make multiple predictions for an instance by identifying associations between labels, thereby enhancing prediction capability. The paper presents a new method for combining improved chicken swarm optimization (ICSO) feature selection with recurrent neural networks (RNNs) for MLC. ICSO overcomes feature selection issues and minimizes noise and redundancy in imbalanced datasets, whereas RNNs learn label dependencies to achieve higher accuracy. It is initially normalized using the Z-score method and then analyzed for dimensionality reduction using principal component analysis (PCA). Our approach is more accurate, more precise, better at recall, and achieves higher F-measure and specificity, and a lower error rate than current methods such as multi-label k-nearest neighbors (ML-KNN), fuzzy rough set learning with label-specific features (FRS-LIFT), and adaptive synthetic data for multi-label classification (ASD-MLC). The paper shows that ICSO is useful for augmenting RNN-based multi-label classification and may be applied in medical diagnosis and bioinformatics.
M. Priyadharshini, Narendruni Lakshmi Priya, Anitha Vippdapu et al.· Bulletin of Electrical Engin...· 0 citations
Multi-class imbalanced datasets are ubiquitous in domains like medical diagnostics, fraud detection, and learning performance classification, where minority classes are critical but underrepresented, and class overlap introduces noise and ambiguous boundaries. Prior research has explored oversampling and undersampling techniques, generative approaches, and hybrid sampling methods to balance multi-class datasets, aiming to enhance classifier performance across imbalanced and overlapping classes. These methods often produce noisy samples near class boundaries, fail to preserve class-specific characteristics, or struggle with high-dimensional data and severe imbalances, particularly in multi-class settings, resulting in poor performance on minority classes. In this article, we propose SafeVAE-GAN, a resampling framework that integrates a conditional β-Variational Autoencoder (β-VAE) with Wasserstein GAN to generate high-quality minority samples. Our framework includes two complementary variants: SafeVAE-GAN
O
, an advanced oversampling method using dual Safe Region Constraints; and SafeVAE-GAN
H
, a hybrid variant that further performs targeted undersampling of majority samples near ambiguous boundaries based on k-NN heterogeneity. Across 26 benchmark datasets, the proposed framework demonstrates competitive performance and consistent improvements in Macro-F1, Macro G-mean, mGM and Macro-MCC over strong baselines. Overall, SafeVAE-GAN provides a robust and scalable solution for multi-class imbalanced learning across diverse domains.
Nhiem Ba Nguyen, Sinh Van Nguyen, B. Nguyễn· Vietnam Journal of Computer...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.