Skip to content

Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models

Sep 2026 · 0 citations · 58 references
Computer Science

TL;DR

This work investigates how fine-tuning alters internal representations in LLMs, including attention patterns and layer-wise activations, and examines whether these changes are linked to task-relevant components identified by EAP that drive task performance.

Abstract

Fine-tuning has emerged as a widely adopted approach for adapting LLMs to a variety of downstream tasks. However, how it reshapes their internal mechanisms remains poorly understood. To address this, we investigate how fine-tuning alters internal representations in LLMs, including attention patterns and layer-wise activations, and examine whether these changes are linked to task-relevant components identified by EAP (e.g., attention heads and logit-level activations) that drive task performance. We find that EAP-identified components are concentrated within specific layers, indicating a degree of functional localisation in how models internalise task-specific behavior. Notably, the distribution of these components across layers is largely uncorrelated with the layers undergoing the most substantial representational changes during fine-tuning. Furthermore, we observe that overlap in EAP-identified components across tasks does not translate into cross-task performance transfer if the tasks are different in nature (e.g. classification vs. generative tasks). More specifically, fine-tuning on one task can lead to a degradation of performance on another when the two tasks exhibit a high degree of overlap in their EAP-identified components.

View source

Similar papers

Delayed Convergence and Emergent CoT Reliance in RL-Tuned Language Models

This analysis reveals that intermediate log-probability is an unreliable indicator for reasoning capability; instead, reasoning performance results from a shift of internal confidence allocation where RL fine-tuning delays internal convergence, exhibiting prolonged mid-layer exploration before converging sharply at the...

Pablo Pérez-Lázaro, Rocío Aznar-Gimeno, F. J. Lacueva-Pérez et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A mechanistic study of language model introspection

Large language models (LLMs) can sometimes report perturbations to their internal activations---even when the input provides no evidence that an intervention occurred. How do models detect and localize such internal changes? We study this question using a controlled task that keeps the input text fixed. We either injec...

Jia-Hong Zou, Xiang-Kun Sun, Ling-Kai Kong et al. · 0 citations
Preprint Aug 2026

Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching

Large Language Models (LLMs) demonstrate remarkable problem-solving capabilities when guided by Chain-of-Thought (CoT) prompting, yet the internal mechanisms underlying these improvements remain poorly understood. In this work, we investigate where CoT-related causal effects emerge across the generated reasoning trajec...

Murat Dura, Serkan Öztürk, Selma Tekir · 1 citation · ⚡1
#natural language process... Preprint Sep 2026

Layer-Informed Fine-Tuning via Three-Stage Functional Segmentation of LLMs

In recent years, the performance of large language models (LLMs) on reasoning tasks has been remarkable, even surpassing human capabilities on various benchmarks. However, there remains a lack of clear understanding in the academic community regarding how the structure and internal parameters of LLMs progressively solv...

Jun-Ning Shao, Si-Wei Wang, Zhi-Xuan Fang · 0 citations
Preprint Aug 2026

Architecture-Dependent Causal Transfer of Activation States Across Large Language Models

End-to-end activation-state transfer between LLMs, as currently implemented, is architecture-dependent rather than universal, and it is concluded that end-to-end activation-state transfer between LLMs is architecture-dependent rather than universal.

Fernando Cardenas Piepereit · 3 citations
Aug 2026

GLA-LoRA: Parameter-efficient LLM fine-tuning with global-local knowledge alignment.

GLA-LoRA establishes a unified learning strategy that synergistically integrates multi-granular contrastive learning with knowledge distillation and establishes that explicit global-local knowledge alignment is essential for achieving high-fidelity, parameter-efficient fine-tuning across diverse language tasks.

Hao Wu, Jian-Qi Gao, Xiangfeng Luo · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.