Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning
An attack is proposed, FAB (Finetuning-activated Adversarial Behaviors), which compromises an LLM via meta-learning techniques that simulate downstream finetuning, explicitly optimizing for the emergence of adversarial behaviors in the finetuned models.