Skip to content
Preprint

From Exposure to Expectation: Frequency, Surprisal, and Language Across Development in Spanish

Aug 2026 · 0 citations · 24 references
Computer Science

Abstract

Surprisal, the negative log-probability a language model assigns to a word given its preceding context, reliably predicts adult reading times. Does it contribute as much to explaining when children acquire individual words? Frequency reflects a learner's cumulative exposure to a word, whereas surprisal reflects how predictable a single occurrence is given its context. We investigate this question across two corpus-based studies of Spanish. In Study 1, we modeled age of acquisition (AoA) for 225 Spanish nouns using lexical frequency and contextual diversity from child-directed speech, plus surprisal from three language models differing in architecture and training language (BETO, BERTIN, mGPT). Frequency strongly predicted AoA (r=-.597, p<.001); surprisal added little beyond frequency and word length, including in a naturalistic-context analysis. In Study 2, we modeled adult fixation durations in the Chilean Spanish subsample of the Multilingual Eye-movement Corpus (MECO Wave 2), using mGPT surprisal alongside two independent frequency measures. Surprisal robustly predicted longer fixation durations after controlling for frequency and word length, consistent across both frequency sources. A matched word-type-level comparison showed the surprisal-behavior association was stronger in reading than in acquisition (z=3.63, p<.001). The findings suggest cumulative lexical exposure and contextual predictability play different roles across the language trajectory: frequency is particularly informative about when early lexical representations are acquired, whereas surprisal captures moment-to-moment processing difficulty in an already-established linguistic system. We discuss this pattern in relation to usage-based and entrenchment-based accounts of lexical development and to the evaluation of language models as models of human language behavior.

View source

Similar papers

Open access Aug 2026

Cognitive Processes of Probabilistic Prediction in Reading: Language Model Surprisal Across Model Sizes, Token Granularity, and Reading Paradigms

Surprisal, the negative log probability of a word given its context, is the dominant computational metric for quantifying reading difficulty and a common item difficulty estimator in reading research. Yet how the language model family, surprisal granularity, and corpus type jointly shape the surprisal–reading time link remains unclear. We conducted a secondary analysis of two public English reading corpora: the Natural Stories Corpus (self-paced reading, 181 readers) and the Provo Corpus (eye tracking with cloze norms). Surprisal was computed from GPT 2 (Small, XL), Llama 2 (7B, 13B), and Llama 3 (8B), together with a 5-gram baseline and human cloze norms. Linear mixed-effects models tested the baseline contribution of neural surprisal, inverse scaling within model families, word-level versus sub-word aggregation, and the linking function shape. Neural surprisal contributed reliable variance above a strong baseline in both corpora. A clear inverse scaling pattern emerged: GPT 2 Small produced the largest fit improvements, exceeding GPT 2 XL, Llama 2 13B, and Llama 3 8B. Word-level aggregation outperformed sub-word aggregation, especially for measures of early lexical access. Non-parametric analyses supported an approximately linear linking function, and cloze norms carried information not fully captured by neural surprisal. These findings show that larger models are not automatically better cognitive models of reading and that surprisal granularity is not a neutral analytic choice.

Shu-Ting Liu, Yong Mei · 0 citations
Open access Jul 2026

Predictive Sentence Parsing in Monolingual and Bilingual Learners.

There is now substantial empirical support for incremental models of speech processing where language is processed in real-time as it unfolds. Developmental studies indicate that very young children process speech incrementally, demonstrating the ability to "listen ahead." The predictive use of linguistic information to anticipate upcoming words has primarily been studied in monolingual children from widely represented linguistic communities (e.g., English speakers). In this study, bilingual Mandarin-English-learning and monolingual English-learning Singaporean children were tested on their use of verb information to predict upcoming food and body part nouns, including in the presence of mispronunciations of the same nouns. When nouns were preceded by a constraining verb, children's target fixation increased. However, overall target fixation remained higher for correct pronunciations than for mispronunciations, suggesting that children were also sensitive to mispronunciations. Further, constraining verbs facilitated children's target fixation even in the presence of speech errors. The difference between target fixation on correct pronunciation compared to mispronunciation trials was smaller in bilinguals than monolinguals when verbs were present, and this effect appeared at a later time window across trials for monolinguals than it did for bilinguals, suggesting earlier resolution of conflicting linguistic cues in bilinguals. Findings contribute to growing knowledge about how language experience interacts with the temporal dynamics of incremental sentence processing in young learners.

S. J. Rajendra, Qiqi Cheng, Leher Singh · 0 citations
Preprint Jul 2026

A scaling law of contextual persistence in human language

Human language exhibits lawful structure at the level of words (frequency, vocabulary growth) and word pairs (co-occurrence across distance). Here we show that the arrangement of words in sequence -- a central determinant of meaning -- obeys a comparable law. Using large language models as probabilistic probes, we measured the reduction in target perplexity conferred by prior context at distance d beyond that of the same words scrambled; this difference, the contextual persistence function P(d), isolates the influence of arrangement. Across ten corpora spanning six language families and written and spoken modalities, P(d) decayed approximately as 1/d ($P(d) \propto d^{-\alpha}$, mean $\alpha = 1.04$; median $r^2 = 0.96$). The effect vanished in scrambled and synthetic controls, replicated across independent probes, and did not appear in genomic or protein sequences under domain-native models. An exponent near 1 distributes contextual influence approximately uniformly across logarithmic timescales. The results establish a scaling law of contextual persistence in human language.

E. Barenholtz · 0 citations
Open access Jul 2026

Word Predictability as a Measure of Second Language Proficiency

This study introduces predictability BERT , a novel metric for assessing second language (L2) proficiency based on the predictability of word choices in learner language production. Using BERT (Devlin et al., 2019), we calculated the conditional probability of each word in a text given its surrounding context. We evaluated predictability BERT on two datasets: the Lexical Proficiency Corpus ( N  = 480), containing analytic ratings of lexical proficiency, and TOEFL 11 ( N  = 11,000; Blanchard et al., 2013), containing standardized language proficiency scores. Results show that predictability BERT correlated with ratings of collocation accuracy ( r  = .80) and lexical proficiency ( r  = .73). In a multilevel model, predictability BERT was the strongest predictor of language proficiency compared to conventional measures of lexical and phraseological sophistication, explaining 59% of variance in TOEFL scores. These findings suggest high‐proficiency L2 learners make more predictable word choices, supporting usage‐based theories emphasizing the role of statistical learning in L2 development.

Langdon Holmes, Scott Crossley, Wesley Morris et al. · 0 citations
Review Open access 2026

Large Language Models as Distributional Baselines for Language Tasks

The central contributions of this paper articulate the conditions under which distributional predictability threatens the internal validity of an experiment and provide concrete recommendations for how to control for this potential confound.

Sean Trott, James A. Michaelov, Cameron R. Jones et al. · 0 citations
Open access Aug 2026

Attention or prediction? Characterizing the top-down influence of predictive context on speech encoding

Theories of predictive coding propose that perception is the process of inferring the causes of our sensory input by comparing that input with predictions derived from our internal models of the world. Such predictive processes are thought to play a central role in language comprehension, however, robust neurophysiological evidence for such processes, particularly during natural speech perception, remains limited. Previous work has suggested that the early auditory encoding of words in natural speech is influenced by their preceding linguistic context. However, it remains unclear whether this effect is driven by prediction per se or dynamic modulations of attention based on contextual uncertainty. To distinguish between these alternatives, we recorded electroencephalography from 17 healthy adults while they listened to slightly changed audiobook. Specifically, we identified and replaced several unsurprising content words with more surprising words. We quantified the early auditory encoding of words using speech-envelope reconstruction accuracy within 100-ms time window after word onset and examined its relationship to word surprisal and contextual uncertainty. We found that more surprising words showed enhanced early auditory encoding despite matched contextual constraint. Moreover, the temporal profile of this enhancement depended on when the incoming speech signal diverged from the predicted phonological sequence, consistent with the emergence of prediction-error responses. Linear mixed-effects modeling further revealed that word surprisal had a substantially stronger influence on early auditory encoding than contextual uncertainty. Together, these findings indicate that the early auditory encoding of words during naturalistic speech perception is more strongly associated with predictive computations than with uncertainty-driven attentional gain. Significance Statement During natural speech comprehension, contextual information influences how the brain processes incoming sensory input. However, whether this context-dependent modulation of the early auditory encoding of words reflects predictive computations or dynamic changes in attentional gain has remained unresolved. By combining a naturalistic speech paradigm with a stimulus manipulation that varies word surprisal while controlling contextual uncertainty, we show that the early auditory encoding of words is driven by word surprisal under matched contextual constraint. Moreover, the temporal dynamics of this modulation closely follow the point at which the incoming speech signal departs from the predicted phonological sequence. These findings provide neurophysiological evidence that predictive computations contribute to the context- dependent modulation of early auditory processing during natural speech comprehension.

Alyssa Horng, Wei-Ching Lin, Ilaria Benciolini et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.