Skip to content
Open access

Cognitive Processes of Probabilistic Prediction in Reading: Language Model Surprisal Across Model Sizes, Token Granularity, and Reading Paradigms

Aug 2026 · Journal of Intelligence · Vol 14 · 0 citations · 35 references
Medicine

Abstract

Surprisal, the negative log probability of a word given its context, is the dominant computational metric for quantifying reading difficulty and a common item difficulty estimator in reading research. Yet how the language model family, surprisal granularity, and corpus type jointly shape the surprisal–reading time link remains unclear. We conducted a secondary analysis of two public English reading corpora: the Natural Stories Corpus (self-paced reading, 181 readers) and the Provo Corpus (eye tracking with cloze norms). Surprisal was computed from GPT 2 (Small, XL), Llama 2 (7B, 13B), and Llama 3 (8B), together with a 5-gram baseline and human cloze norms. Linear mixed-effects models tested the baseline contribution of neural surprisal, inverse scaling within model families, word-level versus sub-word aggregation, and the linking function shape. Neural surprisal contributed reliable variance above a strong baseline in both corpora. A clear inverse scaling pattern emerged: GPT 2 Small produced the largest fit improvements, exceeding GPT 2 XL, Llama 2 13B, and Llama 3 8B. Word-level aggregation outperformed sub-word aggregation, especially for measures of early lexical access. Non-parametric analyses supported an approximately linear linking function, and cloze norms carried information not fully captured by neural surprisal. These findings show that larger models are not automatically better cognitive models of reading and that surprisal granularity is not a neutral analytic choice.

Read PDF

Similar papers

Preprint Aug 2026

From Exposure to Expectation: Frequency, Surprisal, and Language Across Development in Spanish

Surprisal, the negative log-probability a language model assigns to a word given its preceding context, reliably predicts adult reading times. Does it contribute as much to explaining when children acquire individual words? Frequency reflects a learner's cumulative exposure to a word, whereas surprisal reflects how predictable a single occurrence is given its context. We investigate this question across two corpus-based studies of Spanish. In Study 1, we modeled age of acquisition (AoA) for 225 Spanish nouns using lexical frequency and contextual diversity from child-directed speech, plus surprisal from three language models differing in architecture and training language (BETO, BERTIN, mGPT). Frequency strongly predicted AoA (r=-.597, p<.001); surprisal added little beyond frequency and word length, including in a naturalistic-context analysis. In Study 2, we modeled adult fixation durations in the Chilean Spanish subsample of the Multilingual Eye-movement Corpus (MECO Wave 2), using mGPT surprisal alongside two independent frequency measures. Surprisal robustly predicted longer fixation durations after controlling for frequency and word length, consistent across both frequency sources. A matched word-type-level comparison showed the surprisal-behavior association was stronger in reading than in acquisition (z=3.63, p<.001). The findings suggest cumulative lexical exposure and contextual predictability play different roles across the language trajectory: frequency is particularly informative about when early lexical representations are acquired, whereas surprisal captures moment-to-moment processing difficulty in an already-established linguistic system. We discuss this pattern in relation to usage-based and entrenchment-based accounts of lexical development and to the evaluation of language models as models of human language behavior.

Francisco López · 0 citations
Jul 2026

Predictability effects in natural reading are logarithmic: Evidence from an eye-movement replication of Brothers and Kuperberg (2021).

The question of whether the relationship between a word's predictability and its processing time is linear or logarithmic is of substantial theoretical importance, as it arbitrates between theories of the predictability effect (preactivation vs. surprisal) and has implications for sentence processing generally. While several previous corpus studies have obtained evidence for a logarithmic relationship, Brothers and Kuperberg (2021b) obtained evidence for a linear relationship in a large self-paced reading study with well-controlled experimental materials. Here, the authors use Brothers and Kuperberg's materials in an eye tracking during reading experiment. The authors find clear evidence for a logarithmic relationship between predictability and the eye-movement measures of first fixation duration and gaze duration; this relationship is clearest when using predictability estimates from the large language model Generative Pre-Trained Transformer-2, which can distinguish small differences in predictability at the low end of the scale. The authors find that this conclusion is robust to log transformation of the reading time measures, and the authors find that the relationship between predictability and the log odds of word skipping may also be logarithmic. These results support surprisal as an account of predictability effects in natural reading. (PsycInfo Database Record (c) 2026 APA, all rights reserved).

Ryan Buggy, Stephanie Cho, Cory Shain et al. · 1 citation
Open access Aug 2026

Attention or prediction? Characterizing the top-down influence of predictive context on speech encoding

Theories of predictive coding propose that perception is the process of inferring the causes of our sensory input by comparing that input with predictions derived from our internal models of the world. Such predictive processes are thought to play a central role in language comprehension, however, robust neurophysiological evidence for such processes, particularly during natural speech perception, remains limited. Previous work has suggested that the early auditory encoding of words in natural speech is influenced by their preceding linguistic context. However, it remains unclear whether this effect is driven by prediction per se or dynamic modulations of attention based on contextual uncertainty. To distinguish between these alternatives, we recorded electroencephalography from 17 healthy adults while they listened to slightly changed audiobook. Specifically, we identified and replaced several unsurprising content words with more surprising words. We quantified the early auditory encoding of words using speech-envelope reconstruction accuracy within 100-ms time window after word onset and examined its relationship to word surprisal and contextual uncertainty. We found that more surprising words showed enhanced early auditory encoding despite matched contextual constraint. Moreover, the temporal profile of this enhancement depended on when the incoming speech signal diverged from the predicted phonological sequence, consistent with the emergence of prediction-error responses. Linear mixed-effects modeling further revealed that word surprisal had a substantially stronger influence on early auditory encoding than contextual uncertainty. Together, these findings indicate that the early auditory encoding of words during naturalistic speech perception is more strongly associated with predictive computations than with uncertainty-driven attentional gain. Significance Statement During natural speech comprehension, contextual information influences how the brain processes incoming sensory input. However, whether this context-dependent modulation of the early auditory encoding of words reflects predictive computations or dynamic changes in attentional gain has remained unresolved. By combining a naturalistic speech paradigm with a stimulus manipulation that varies word surprisal while controlling contextual uncertainty, we show that the early auditory encoding of words is driven by word surprisal under matched contextual constraint. Moreover, the temporal dynamics of this modulation closely follow the point at which the incoming speech signal departs from the predicted phonological sequence. These findings provide neurophysiological evidence that predictive computations contribute to the context- dependent modulation of early auditory processing during natural speech comprehension.

Alyssa Horng, Wei-Ching Lin, Ilaria Benciolini et al. · 0 citations
Open access Aug 2026

Brain–Language Alignment During Naturalistic Reading and Its Disruption by Mind-Wandering

Encoding models offer a principled framework for linking computational representations of language to neural activity, but most electroencephalography (EEG) evidence for brain–language alignment comes from tightly controlled, word-by-word reading paradigms. Whether such alignment is detectable during naturalistic reading, and how it is affected by lapses in attention, remains unclear. We addressed these questions using ROAMM, a multimodal dataset containing simultaneous EEG and eye-tracking recordings with time-resolved mind-wandering (MW) annotations from 44 participants reading naturalistic texts. Ridge regression encoding models were trained to predict fixation-aligned EEG spectral power and fixation-related potentials (FRPs) from five word-embedding models (GloVe, word2vec, BERT, GPT-2, and Llama 3). Using permutation testing with false discovery rate correction, we found statistically reliable brain–language alignment across both feature types, with contextual embeddings outperforming static embeddings. Spectral alignment was strongest in the alpha and low-beta bands over parietal electrodes, while FRP-based alignment peaked 200–300 ms after fixation onset over central and parietal-occipital regions. Leveraging ROAMM’s span-level MW annotations, we further show that brain–language alignment is systematically reduced during MW, an effect that was substantially larger for oscillatory (PSD) than for event-related (FRP) features. These findings demonstrate that modern language-model representations are reflected in EEG activity during naturalistic reading despite the modality’s inherent noise, and that fluctuations in attention constitute an underappreciated source of variability in brain–language encoding studies.

Haorui Sun, D. Jangraw · 0 citations
Open access Aug 2026

A Predictive Comparison of the Selective Attention Hypothesis and Contextual Priming Theory in Processing English It-Extraposition: An Eye-Tracking Study

When readers encounter an English it-extraposition structure (e.g., It is obvious that the minister resigned), two fundamentally different processing mechanisms could guide the resolution of the cataphoric dependency. The Selective Attention Hypothesis (SAH) holds that expletive it concentrates attentional resources on the matrix predicate, creating an integration cost that spills over onto the extraposed clause. Contextual Priming Theory (CPT), by contrast, claims that the matrix context pre-activates features of the upcoming clause, so that higher predictability reduces downstream effort. We adjudicated between these predictions in the first eye-tracking study of it-extraposition with advanced Libyan learners of English (N = 44). Participants read extraposed and non-extraposed sentences while their eye movements were recorded. Two linear mixed-effects models were constructed: an SAH model incorporating log frequency of the matrix adjective, and a CPT model incorporating offline cloze probability of the embedded-clause verb. The CPT model provided a significantly better fit than a baseline model (Δ − 2LL = 6.52, p = .011), with each one-unit increase in centred predictability reducing total reading time by 0.35 ms (f² = 0.04). The SAH predictor did not approach significance (Δ − 2LL = 1.23, p = .267). Follow-up analyses of early measures, regressions, and adjective-frequency interactions offered no support for an attentional bottleneck. We discuss the results in light of surprisal theory, Bayesian predictive coding, and constraint-based processing, and consider the possibility that SAH effects if present may require larger samples or different operationalisations to detect. The findings confirm that advanced L2 learners exploit the predictive affordances of the matrix context, and they carry direct implications for teaching grammar and reading in the Libyan EFL classroom.

Muwahib Mohammed, Abdul Salam, Al-Shabouki · 0 citations
Preprint Jul 2026

Encoding EEG Signals to Examine Human-Like Next-Word Prediction Behaviour in Language Models

Language models (LMs) are trained to excel at predicting the next word in the sequence given prior context, and humans also share this predictability in reading comprehension. Neuroscience research reveals that next-word predictability influences brain response, as recorded at millisecond resolution using electroencephalography (EEG). While our evidence indicates that advanced LMs achieve accuracies closely aligned with human performance at the next-word prediction task, this raises the question: Does higher prediction accuracy necessarily mean that these models adequately capture the cognitive signals associated with human reading comprehension? Here, we generate regressors for both humans and LMs based on two information measures, including top-1 prediction and surprisal, to predict event-related potential (ERP) elicited from EEG recordings which reflect different stages of cognitive processing during reading. We argue that modelling ERP patterns offers fine-grained analysis of the cognitive plausibility of various LMs during reading. Our results indicate that only surprisal potentially correlates with language-processing ERPs, especially for open-class words with high semantic content. Moreover, our findings challenge the assumption that scaling LMs with increased parameters and computational budgets will consistently lead to improved convergence with human-like linguistic processing.

B. M. Quach, B. Nguyen, C. Gurrin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.