ASCENT: First-Order Optimal Fine-Tuning with Recalibration for Safety--Utility Co-Enhancement
Supervised fine-tuning can substantially improve the downstream utility of large language models (LLMs) but may compromise their safety. Existing safety-preserving methods constrain downstream updates using safety-related parameters or subspaces, but mainly focus on safety preservation rather than joint safety and util...