Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models
This work investigates how fine-tuning alters internal representations in LLMs, including attention patterns and layer-wise activations, and examines whether these changes are linked to task-relevant components identified by EAP that drive task performance.
Ling-Fang Li, Procheta Sen, Shubham Das et al.
· 0 citations