Jul 2026· Journal of Advanced Computational Intelligence and Intelligent Informatics· Vol 30, pp. 1243-1257· 0 citations· 17 references
Computer Science
TL;DR
Comparing input embeddings and hidden states from multiple language model families within both encoding frameworks, which map text features to brain responses, and decoding frameworks, which reconstruct linguistic features from brain activity shows that input embeddings offer a strong and interpretable baseline for representational alignment between language models and brain activity.
Abstract
Systematically comparing how linguistic representations relate to brain activity has become an important topic in the field of computational neuroscience. Prior studies have mainly relied on contextual hidden states combined with linear regression, leaving open questions about the role of static input embeddings and the benefits of nonlinear mappings. In this study, we compare input embeddings and hidden states from multiple language model families (BERT, GPT-2, and LLaMA) within both encoding frameworks, which map text features to brain responses, and decoding frameworks, which reconstruct linguistic features from brain activity. We benchmarked voxel-wise ridge regression against bidirectional long short-term memory (BiLSTMs) models, using repeat-split cross-validation and explainable variance normalization on functional magnetic resonance imaging (fMRI) data from three subjects. Our analyses demonstrate that input embeddings, despite being context-invariant, remain competitive and, in some cases, outperform hidden states, while BiLSTMs provide modest but region-specific improvements over ridge regression. Fine-grained voxel-level results further revealed distinct cortical distributions of stable versus context-dependent features. Together, these findings clarify the trade-off between predictive performance and interpretability and highlight that input embeddings offer a strong and interpretable baseline for representational alignment between language models and brain activity.
Foundation-model features are increasingly used to ask what information neural activity represents, often by comparing prediction gains between nested encoding models. We show that such multimodal contrasts can change sign when only the conditioning predictor is reconstructed. Using fMRI from the Natural Scenes Dataset...
Lucas Nadolskis, Galen Pogoncheff, Michael Beyeler· 0 citations
Encoding models offer a principled framework for linking computational representations of language to neural activity, but most electroencephalography (EEG) evidence for brain–language alignment comes from tightly controlled, word-by-word reading paradigms. Whether such alignment is detectable during naturalistic readi...
Foundation models pre-trained on large-scale fMRI datasets have shown strong downstream performance, but at substantial data and computation cost. To investigate how much fMRI-specific pre-training is actually needed for such performance, we introduce FReD, which derives fMRI representations from a frozen Deep Compress...
Juhyeon Park, Yeonwook Kim, P. Y. Kim et al.· 0 citations
The brain processes information across distributed circuits, yet a typical experiment records only a few regions, leaving the rest unobserved. Connectivity and latent-embedding methods relate brain regions but do not return the waveform of an unrecorded one. Here we introduce NeuroGate, a framework for cross-regional n...
Ali Zareh, Rufeyda Yağcı, M. K. Özdemir· bioRxiv· 0 citations
Human language processing can be studied through both behavior and brain activity, yet it remains unclear whether these two data types reflect sensitivity to the same information. One influential view holds that both behavioral and neural responses are largely determined by processing effort, often estimated by word su...
Andrea Gregor de Varda, Yevgeni Berzak, Evelina Fedorenko et al.· bioRxiv· 0 citations
Understanding how the brain parses actions and events from time-varying natural inputs is a central challenge in neuroscience. Recent work has used deep neural network (DNN) models to build stimulus-computable fMRI encoding models that predict single-voxel responses to complex natural videos. However, the majority of v...
Iishaan Inabathini, Margaret M. Henderson· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.