Skip to content
Open access

АДАПТИВНА РЕГУЛЯРИЗАЦІЯ ЗНАНЬ ДЛЯ ТРАНСФОРМЕРНИХ АРХІТЕКТУР У ЗАДАЧАХ ПОСЛІДОВНОГО НАВЧАННЯ

Aug 2026 · RADIOELECTRONIC AND COMPUTER SYSTEMS · 0 citations

Abstract

The subject matter of the article is development of latent representation regularization mechanism in transformer-based architecture under conditions of continuous learning with domain shifts. Modern language models achieve high quality in static learning scenarios, but they remain limited in long-term operation cases, incremental adaptation to new domains and lacking resistance to catastrophic forgetting, especially in the absence of access to previously observed data. This paper explores the possibility of overcoming these limitations by combining uncertainty inspired regularization and forgetting attention mechanisms into one transformer architecture. The goal of the study is to design, implement and validate transformer-based architecture with multi-level representation regularization mechanism that can help transformer-based language models efficiently adapt to alternative data distribution while retaining previously acquired knowledge. The proposed approach aims to achieve an equilibrium between model adaptability and stability in continual learning without requiring complete model retraining or legacy data retention. The tasks to be solved in this study include: formalization of an unified conceptual latent representations regularization method that combines Bayesian uncertainty inspired latent representations regularization with head-wise attention scaling in attention mechanism; implementation of this method into the transformer model; creating the experimental case of continuous language modeling with sequential domain shifts; give quantitative estimation of model forgetting and stability in prediction quality; compare the proposed model to classical naive fine-tuning, LoRA and parameter regularization methods. The conclusions demonstrate that the proposed method achieves lower forgetting, lower perplexity on previously learned domains and a better stability–plasticity trade-off than naive fine-tuning, LoRA and Elastic Weight Consolidation, while requiring comparable computational resources. The scientific novelty of proposed approach consists in development of layer-selective latent regularization framework for continual language modeling which integrates an attention with forgetting mechanism with preserving domain-invariant representations through statistical alignment in latent space for reducing forgetting in continual learning scenarios. Unlike existing approaches to continual learning that consider model regularization either on parameter or memory level (by using previous data), the proposed approach moves regularization into representation space, where it uses both direct regularization via proposed uncertainty-based regularization and indirect regularization via attention with forget gate, ensuring the models’ possibility of stable continual language modeling in non-stationary environments.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.