The central contributions of this paper articulate the conditions under which distributional predictability threatens the internal validity of an experiment and provide concrete recommendations for how to control for this potential confound.
Abstract
In order to ask questions about the mechanisms underpinning human cognition, researchers must control for properties of stimuli that could confound detected effects. In experiments involving linguistic stimuli, this includes properties like frequency, length, and neighborhood size of those stimuli, which are known to affect behavioral and neural responses. With improvements in the performance and usability of language models, it is now possible to also control for how predictable stimuli and their parts are, on the basis of the distributions of words alone: their distributional predictability. This coincides with a resurgence of interest in the possibility that statistical language learning may underlie a broad range of human cognitive phenomena; indeed, there are both theoretical and empirical reasons to believe that humans rely on distributional information during certain cognitive tasks. This creates a confound, whereby experimental operationalizations of psychological constructs with linguistic stimuli may not in fact be testing what they are intended to test. Thus, the central contributions of this paper are twofold: first, we articulate the conditions under which distributional predictability threatens the internal validity of an experiment; and second, we provide concrete recommendations for how to control for this potential confound. Beyond these primary contributions, we survey techniques for measuring distributional predictability, review theoretical and empirical work supporting the role of distributional statistics in human cognition, and present several case studies illustrating the range of possible outcomes—from the “distributional baselines” only marginally affecting theoretical inferences to constituting fully deflationary confounds. We also enumerate and address potential objections to this approach. This paper is primarily intended for researchers in psychology, cognitive science, and linguistics who use linguistic stimuli but have not yet incorporated distributional baselines into their work.
Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. For in-distribution simulations, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choice history -- across 27 experiments, and permute trial order. Masking the content of stimuli and feedback destroys 75.7% of learned information and pushes models below chance, demonstrating that choice history alone does not account for performance. Permutation reveals invariance on tasks with independent trials but sensitivity where trial order is determined by prior responses. Small cognitively fine-tuned models therefore show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by the paradigms seen in training.
Human language processing can be studied through both behavior and brain activity, yet it remains unclear whether these two data types reflect sensitivity to the same information. One influential view holds that both behavioral and neural responses are largely determined by processing effort, often estimated by word surprisal together with the context-independent properties of word frequency and length. At the same time, neural responses have been shown to encode richer aspects of linguistic content, including meaning. Here, we use neural network language models to operationalize these alternatives and systematically compare, within the same analytic computational framework, the predictive power of low-dimensional effort-based predictors and high-dimensional embedding representations that encode contextualized linguistic content, including meaning. Across 8 behavioral datasets and 5 neural datasets (4 fMRI and 1 ERP), we find that processing effort captures substantial variance in both behavioral and neural measures of language processing, in line with much previous work. However, for brain responses—but not for behavioral measures—embedding representations carry substantial predictive power beyond the estimates of processing effort. These results therefore suggest that neural data provide access to rich, high-dimensional dynamics of language comprehension, whereas behavioral data reflect a bottlenecking of these dynamics into a small set of theoretically motivated properties of contextualized linguistic input. Significance Statement Two research communities study language comprehension as it unfolds in real time: psycholinguists use behavioral measures, such as eye movements during reading, and neuroscientists measure brain activity. The two are rarely studied together, but evidence from both must be integrated into a unified theory of language processing. Here we analyze both brain and behavioral responses within a single framework based on language models, comparing two long-standing accounts of what drives responses to language: processing effort versus meaning and other features not reducible to effort. We find that behavior is dominated by effort, whereas brain responses also reflect meaning. Developing a unified theory requires both kinds of data, but with a clear understanding of which levels of representation each measure reflects.
Andrea Gregor de Varda, Yevgeni Berzak, Evelina Fedorenko et al.· bioRxiv· 0 citations
Linguistic typology has identified properties shared by the world’s languages, as well as features with respect to which languages diverge, including infrequent or rare phenomena. Cognitive biases are one important source of language universals, through their indirect effect on language change via their impact on language learning and language use. At the same time, typologically rare phenomena can exist only if the human brain is able to accommodate them. Nonetheless, our knowledge of how we process language mostly derives from the analysis of a limited set of languages associated with WEIRD (Western Educated Industrialized Rich Democratic) societies. Including non-WEIRD languages in our psycholinguistic experiments stands to benefit the cognitive science of language and, in addition, linguistic typology, given that languages do best what speakers do most, which in turn depends on what our brains do most. In this paper, the benefits of a less biased approach to the cognitive science of language diversity are illustrated through the consideration of a research project aimed at determining the impact of the typological diversity of human languages on different memory types, specifically, procedural vs. declarative memory, with the latter encompassing two subtypes of relevance to linguistic structure, semantic and episodic. Our focus is on typological features that have been shown, in the past, to correlate with sociolinguistic factors. Overall, we aim at delving into an array of multi-directional feedback relationships among memory types, sociolinguistic properties, and typological features.
Antonio Benítez-Burraco, Sihan Chen, D. Gil· Linguistic Typology at the C...· 0 citations
Statistical learning (SL) - the ability to detect patterns in sensory input without explicit instruction - is crucial for building internal models of the environment. In humans, it notably supports language acquisition, including word segmentation and grammar learning. Evidence across primates, songbirds, rodents, and insects indicate that SL is a widely shared evolutionarily conserved ability. The computational complexity of the mechanisms involved; however, varies between species, likely reflecting specific cognitive limitations. Alternatively, specific competences may have evolved to support the emergence of demanding, ecologically relevant, functions such as complex communication systems. These findings challenge the idea of a human-specific SL module while raising key questions about its evolutionary drivers and underlying mechanisms. Addressing these questions requires the expansion of cross-species, cross-modal studies with ecologically valid, unsupervised paradigms. Comparative approaches are indeed essential to uncover shared properties and species-specific SL adaptations. Framing SL as a foundational component of cognition, informed by animal research, offers new insights into brain function, and implicit learning in light of the evolution of sophisticated, emergent cognitive abilities such as complex communication systems.
Laure Tosatto, A. Avarguès-Weber· Current Opinion in Neurobiol...· 0 citations
These findings demonstrate that LMs can recover human-like linguistic generalizations from impoverished input and provide a controlled framework for investigating the mechanisms underlying such biases.