Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning
Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence. However, this adaptation process can also degrade alignment properties that were present in the source model. Recent work has shown that large language models can b...