Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning
An attack is proposed, FAB (Finetuning-activated Adversarial Behaviors), which compromises an LLM via meta-learning techniques that simulate downstream finetuning, explicitly optimizing for the emergence of adversarial behaviors in the finetuned models.
Thibaud Gloaguen, Mark Vero, Robin Staab et al.
· 4 citations