Skip to content
Open access

DMAE: dual-memory active ensemble learning for multi-class imbalanced concept-drifting data streams

Aug 2026 · Journal of King Saud University: Computer and Information Sciences · Vol 38 · 0 citations · 45 references

Abstract

Multiclass imbalance and concept drift often coexist in real-world data streams, and the challenge becomes more severe under limited labeling budgets, where existing online ensemble and active learning methods still struggle to preserve minority-class recognition while adapting to evolving concepts. To address this issue, we propose DMAE, a dual-memory active ensemble learning method for multiclass imbalanced concept-drifting data streams. DMAE integrates an ensemble classifier, a drift detector, an instance sliding window, a label sliding window, a long-term prototype memory, and an initialization training sequence into a unified online framework for prediction, querying, memory maintenance, and model updating. To improve label efficiency, we develop a variable-threshold uncertainty strategy based on a decomposable asymmetric margin-threshold matrix, which combines a global threshold with class-pair-level corrections to focus queries on genuinely ambiguous, minority-relevant regions while controlling annotation cost. We further design drift-strength-aware dual-memory initialization (SDMI) and a probationary soft-replacement strategy (PDSR), which adaptively balance recent information and long-term prototypes according to online-estimated drift strength, supporting both historical-pattern retention and rapid adaptation under abrupt, gradual, and mixed drifts. We also propose a composite sample-weighting formulation, PCN-Weight, which jointly models boundary difficulty, class-imbalance status, sample–prototype relations, and temporal decay to guide incremental training and strengthen minority-class and hard-region representations. Experiments on 15 synthetic data streams and 5 real-world imbalanced data streams show that DMAE achieves more stable overall performance than ten state-of-the-art ensemble baselines in Accuracy, Kappa, G-Mean, and Recall, while remaining robust under different labeling rates and key hyperparameter settings.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.