Sep 2026· Machine Learning and Knowledge Extraction· Vol 8, pp. 282· 0 citations· 25 references
TL;DR
This work identifies ratio-invariant safety—non-degradation at any imbalance ratio, satisfied by vanilla BCE and q-calculus but violated by all reweighting methods—and gives recommendations spanning moderate to extreme imbalance.
Abstract
Training under extreme class imbalance (>1:100) remains an open problem in weakly supervised learning. The standard remedy—loss-level reweighting (focal loss, asymmetric loss, class-balanced loss)—is widely adopted, yet its behavior at extreme ratios in Multiple Instance Learning (MIL) is poorly understood. On digital breast tomosynthesis (attention-based pooling over frozen EfficientNet-B3 features), we study the optimization bounds under extreme bag-level MIL imbalance (1:251), intervening at two levels: the loss surface (via reweighting) and the gradient dynamics (via a novel q-calculus gradient modification using the Jackson q-derivative). All three reweighting strategies degrade classification relative to unweighted binary cross-entropy (BCE), monotonically, eliminating the loss surface as the bottleneck. Extended evaluation (n=20 seeds) shows q-calculus gradient smoothing matches vanilla BCE (p=0.632, Cohen’s d=0.003) despite provably reducing gradient variance, establishing an empirical ceiling on the optimization-achievable area under the precision-recall curve (AUPRC) of 0.0912; loss reweighting defines the floor at 0.055. Focal loss is additionally catastrophically miscalibrated (ECE > 0.44 vs. 0.036 for vanilla BCE), a collapse that persists under adaptive binning. In this regime, exceeding the ceiling points to the data-representation level, not the optimizer. We further identify ratio-invariant safety—non-degradation at any imbalance ratio, satisfied by vanilla BCE and q-calculus but violated by all reweighting methods—and give recommendations spanning moderate to extreme imbalance.
A systematic comparative framework that integrates data-level resampling, cost-sensitive learning, and hybrid approaches to evaluate their performance under varying imbalance ratios and noise levels indicates that hybrid approaches consistently outperform standalone methods, achieving the most stable and balanced perfo...
Tamsir Ariyadi, E. Noche, Nisha Pandey et al.· Journal of Data Science· 0 citations
A semi-supervised weighted stacked autoencoder with spectral peak significance constraints (SSWAEF), a nearest-neighbor consistency voting strategy coupled with a reciprocal-class-size sampling mechanism is first developed to construct a high-quality, class-balanced pseudo-labeled dataset.
Ying-Hao Zhao, Xu Yang, Jian Huang et al.· Structural Health Monitoring· 0 citations
Class imbalance is a critical challenge in the classification of tabular data, since it affects the diagnostic capacity of models in domains such as health and finance. This research compares four synthetic data generation paradigms: traditional interpolation (SMOTE-NC), deep generative models (CTGAN and TVAE), and a h...
Jhonatan Esquivel, Christian Humpiri, Jose M. Vega et al.· International Journal of Adv...· 0 citations
SMOG, an adaptive hybrid oversampling framework that integrates the Synthetic Minority Over-sampling Technique with a Conditional Generative Adversarial Network (GAN)-based difficulty-aware learning strategy, highlights the effectiveness of adaptive hybrid generative strategies for intelligent learning on imbalanced da...
Jatinder Kaur, Vimal Parmar, B. K. Rao et al.· Evolutionary Intelligence· 0 citations
Bounded predictive influence and reliability-guided geometry as complementary mechanisms for imbalanced learning with uncertain labels are supported as complementary mechanisms for imbalanced learning with uncertain labels.
M. Akhtar, J. Akarsh, M. Tanveer et al.· 0 citations
Overall, this dissertation provides a unified investigation into data imbalance, data quality, and data scarcity-three core bottlenecks of modern deep learning-and proposes principled solutions that improve robustness, interpretability, and efficiency across both CV and NLP domains.
Jian Sun· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.