Skip to content
Open access

TDG-LoRA: Token-Level Dynamic Gating for Mitigating Catastrophic Forgetting

2026 · IEEE Access · Vol 14, pp. 149792-149802 · 0 citations · 34 references

Abstract

Parameter-efficient fine-tuning (PEFT), particularly Low-Rank Adaptation (LoRA), is widely used to adapt large language models (LLMs) to specialized downstream domains. However, although the pretrained backbone remains frozen, a domain-adapted LoRA branch may interfere with the model’s original representations and predictions when applied to general-domain inputs, resulting in general capability degradation. To address this problem, we propose TDG-LoRA (Token-Level Dynamic Gating for Low-Rank Adaptation), which introduces a lightweight gating network into each adapted Transformer layer to regulate the contribution of a single LoRA branch according to contextualized token representations. During training, sequence-origin domain labels supervise the tokenwise gate outputs, while an asymmetric loss mask prevents general-domain language-modeling losses from directly updating the LoRA parameters. During inference, the learned gate suppresses unnecessary LoRA contributions on general-domain inputs. We evaluate TDG-LoRA on Llama-3.2-1B and Llama-3.1-8B across mathematical and medical domain adaptation settings using five random seeds. Under a predefined equivalence margin of ±1.0 percentage points, TDG-LoRA is statistically equivalent to standard LoRA on GSM8K and MedQA and to the corresponding base model reference on post-adaptation Massive Multitask Language Understanding (MMLU). On Llama-3.1-8B, TDG-LoRA achieves 58.92% accuracy on GSM8K while retaining 65.18% on MMLU, corresponding to a decrease of only 0.05 percentage points from the base model. These results demonstrate a favorable balance between target domain adaptation and general capability retention under the evaluated settings.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.