Skip to content

ADAPTIVE KNOWLEDGE REGULARIZATION FOR CONTINUAL LEARNING IN TRANSFORMER ARCHITECTURES

TL;DR

The conclusions demonstrate that the proposed method achieves lower forgetting, lower perplexity on previously learned domains and a better stability–plasticity trade-off than naive fine-tuning, LoRA and Elastic Weight Consolidation, while requiring comparable computational resources.

View source

Similar papers

Preprint Aug 2026

Geometry of Forgetting: Representation Flux in Continual Learning

This work proposes FlowLess-R, a representation-space regularization method that constrains replay representations relative to stored references while allowing continued learning and introduces representation flux, a geometric measure of sample-level representation displacement across training.

M. A. Kazanskii · 0 citations
Open access Jul 2026

Navigating parameter space: mitigating catastrophic forgetting in continual learning

Investigation of the influence of fully connected FC layer architecture on parameter regularization in the class incremental learning setting using a modified ResNet-18 trained on the CIFAR-10 dataset provides both a novel parameter regularization strategy and new insights into the interaction between network architect...

Henry Huang · 0 citations
Jul 2026

The Art of Not Forgetting A Local Learning Architecture for Continual Learning

The results suggest that the combination of sparse representations, local learning, and persistent memory is a promising direction for continual learning, while motivating further investigation into the respective roles of learning rules, representations, and architectural design in mitigating catastrophic forgetting.

Ashmith Atmuri, Yashaswini Rao Bhogarajula · 0 citations
Preprint Aug 2026

TASSO: TAsk-Specific Subspace Optimization for Continual Learning of Vision-Language Models

TASSO, a new paradigm that efficiently preserves the latent space geometry while ensuring network plasticity, is introduced with two complementary techniques: subspace learning and geometry-aware knowledge distillation.

Chang-Ming Sun, Francesco Barbato, Matteo Caligiuri et al. · 0 citations
Jul 2026

Latent-LoRA: Compact Latent-Space Adapters with Gradient-Free Routing for Continual Learning

This work observes that pooled token embeddings from a frozen LLM embedding layer already separate task distributions throughout the learning sequence, and concludes that a Gaussian mixture model fitted on these embeddings, without any gradient-based training, is sufficient for task-agnostic adapter selection at test t...

Reza Rahimi Azghan, Gautham Krishna Gudur, Giulia Pedrielli et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.