Catastrophic Forgetting in Parameter-Efficient Fine-Tuning of Small Language Models: A Controlled Comparison of LoRA, Bottleneck Adapters, and Full Fine-Tuning under Sequential Task Learning
Abstract Four small language models of 124M–1.1B parameters were trained on the same five-task stream with five configurations: full fine-tuning, LoRA (r=8 and 16), bottleneck adapters, and LoRA r=8 with 2% replay. We evaluate task accuracy, backward transfer, forgetting, forward transfer, and learning plasticity. Acro...