Skip to content
Open access

Mitigating Catastrophic Forgetting in Incremental Learning Using Hybrid Approach: Interleaving Memory Replay and Parameter Regularization for Sequential Text Classification

Sep 2026 · Electronics · 0 citations · 42 references

Abstract

Catastrophic forgetting is a major challenge for deep learning models when they are incrementally trained on a sequence of new data. Reducing this forgetting in image and video data has been the primary research focus, but less attention has been given to textual domains, where discrete token distributions and semantic shifts occur across different topics. Furthermore, standalone strategies proposed for reducing catastrophic forgetting still have room for improvement. To this end, this paper proposes a synergy of stratified memory replay with parameter regularization for a BiLSTM-based incremental learning model to mitigate catastrophic forgetting in sequential text datasets. The stratified replay mechanism replays a small buffer of historical data samples into current training phases to preserve the old data patterns, while the parameter regularization penalizes modifications to the neural network weights crucial to the past tasks. The proposed approach is evaluated in an incremental training pipeline on distinct textual datasets, including software bug reports (Task A), a news dataset (Task B), and emails (Task C). The evaluation results demonstrate that the baseline neural network experiences catastrophic forgetting as its initial dataset (Task A) accuracy drops from 87.53% to 11.75%. The standalone experience replay approach manages to retain Task A accuracy at 80.01%, down from its peak of 86.56%, while for Task B, it achieves 93.97%, down from the peak of 98.37%. The buffer sensitivity analysis indicates the model accuracy improves with increasing replay buffer size. The parameter regularization approach effectively reduces catastrophic forgetting, but it remains less effective for disruptive text distribution sequences, resulting in noticeable forgetting on prior tasks and reduced plasticity on later tasks. Evaluations of this approach show that Task A accuracy is reduced from 87.87% to 71.09% after training on Task C. The proposed hybrid approach reduces forgetting and preserves Task A and Task B accuracies at 83.73% and 95.04%, respectively. Thus, the empirical evaluations demonstrate that the proposed approach is effective and outperforms the standalone experience replay strategy and parameter regularization, limiting the forgetting on the earliest task to just 4.65% compared to 6.55% forgetting of the replay-based method, while allowing enough plasticity for the final task to reach 98.92% accuracy.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.