This work investigates how fine-tuning alters internal representations in LLMs, including attention patterns and layer-wise activations, and examines whether these changes are linked to task-relevant components identified by EAP that drive task performance.
Abstract
Fine-tuning has emerged as a widely adopted approach for adapting LLMs to a variety of downstream tasks. However, how it reshapes their internal mechanisms remains poorly understood. To address this, we investigate how fine-tuning alters internal representations in LLMs, including attention patterns and layer-wise activations, and examine whether these changes are linked to task-relevant components identified by EAP (e.g., attention heads and logit-level activations) that drive task performance. We find that EAP-identified components are concentrated within specific layers, indicating a degree of functional localisation in how models internalise task-specific behavior. Notably, the distribution of these components across layers is largely uncorrelated with the layers undergoing the most substantial representational changes during fine-tuning. Furthermore, we observe that overlap in EAP-identified components across tasks does not translate into cross-task performance transfer if the tasks are different in nature (e.g. classification vs. generative tasks). More specifically, fine-tuning on one task can lead to a degradation of performance on another when the two tasks exhibit a high degree of overlap in their EAP-identified components.
This analysis reveals that intermediate log-probability is an unreliable indicator for reasoning capability; instead, reasoning performance results from a shift of internal confidence allocation where RL fine-tuning delays internal convergence, exhibiting prolonged mid-layer exploration before converging sharply at the...
Pablo Pérez-Lázaro, Rocío Aznar-Gimeno, F. J. Lacueva-Pérez et al.· 0 citations
Large language models (LLMs) can sometimes report perturbations to their internal activations---even when the input provides no evidence that an intervention occurred. How do models detect and localize such internal changes? We study this question using a controlled task that keeps the input text fixed. We either injec...
Jia-Hong Zou, Xiang-Kun Sun, Ling-Kai Kong et al.· 0 citations
Large Language Models (LLMs) demonstrate remarkable problem-solving capabilities when guided by Chain-of-Thought (CoT) prompting, yet the internal mechanisms underlying these improvements remain poorly understood. In this work, we investigate where CoT-related causal effects emerge across the generated reasoning trajec...
Murat Dura, Serkan Öztürk, Selma Tekir· 1 citation· ⚡1
In recent years, the performance of large language models (LLMs) on reasoning tasks has been remarkable, even surpassing human capabilities on various benchmarks. However, there remains a lack of clear understanding in the academic community regarding how the structure and internal parameters of LLMs progressively solv...
End-to-end activation-state transfer between LLMs, as currently implemented, is architecture-dependent rather than universal, and it is concluded that end-to-end activation-state transfer between LLMs is architecture-dependent rather than universal.
GLA-LoRA establishes a unified learning strategy that synergistically integrates multi-granular contrastive learning with knowledge distillation and establishes that explicit global-local knowledge alignment is essential for achieving high-fidelity, parameter-efficient fine-tuning across diverse language tasks.
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.