DARTS: Decoder-Aware Representation Tuning via Surgery for Model Merging
This work proposes Decoder-Aware Representation Tuning via Surgery (DARTS), which employs a novel entropy-weighted L1 loss to upweight correction at high-entropy positions where errors most affect generation quality, and a per-position additive bias that captures position-dependent error without overparameterization.