TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment
This paper proposes TRACE, which simulates harmful SFT trajectories to produce progressively corrupted model states, and optimizes a safety patch simultaneously across these states, and improves the safety rate by up to 77 percentage points over the best baselines, while maintaining comparable utility to the undefended...