Skip to content
Open access

Population-level linguistic conventionality benefits the robustness, learnability, and cognitive efficiency of emergent languages

2026 · Journal of Language Evolution · Vol 11 · 0 citations · 41 references

Abstract

Human languages are widely understood to be conventional: for a given meaning, a population of language users expect a certain form to be used. Given the lack of unconventional natural languages with which to compare, there is little direct empirical evidence which explains why linguistic conventionality matters. We carry out a series of emergent language experiments with artificial agents designed to exhibit the effect of conventionality on robustness, learnability, and cognitive efficiency. By experimentally manipulating the interaction dynamics through which agents align their vocabularies, we create conditions that either promote or prevent population-level convergence, resulting in conventional and unconventional emergent languages while holding communicative success constant. We find that conventional languages are indeed more robust, better learnable, and more efficient, when compared with less conventional languages.

Read PDF

Similar papers

Open access Jul 2026

Linguistic typology and the cognitive science of non-WEIRD societies: The role of memory types

Linguistic typology has identified properties shared by the world’s languages, as well as features with respect to which languages diverge, including infrequent or rare phenomena. Cognitive biases are one important source of language universals, through their indirect effect on language change via their impact on language learning and language use. At the same time, typologically rare phenomena can exist only if the human brain is able to accommodate them. Nonetheless, our knowledge of how we process language mostly derives from the analysis of a limited set of languages associated with WEIRD (Western Educated Industrialized Rich Democratic) societies. Including non-WEIRD languages in our psycholinguistic experiments stands to benefit the cognitive science of language and, in addition, linguistic typology, given that languages do best what speakers do most, which in turn depends on what our brains do most. In this paper, the benefits of a less biased approach to the cognitive science of language diversity are illustrated through the consideration of a research project aimed at determining the impact of the typological diversity of human languages on different memory types, specifically, procedural vs. declarative memory, with the latter encompassing two subtypes of relevance to linguistic structure, semantic and episodic. Our focus is on typological features that have been shown, in the past, to correlate with sociolinguistic factors. Overall, we aim at delving into an array of multi-directional feedback relationships among memory types, sociolinguistic properties, and typological features.

Antonio Benítez-Burraco, Sihan Chen, D. Gil · 0 citations
Preprint Jul 2026

Exposure is Optional: Learning Unlike Coordination in Language Models

Analysis of internal representations indicate that language models process unlike coordination by treating the conjoined elements as belonging to similar structural categories or through a mechanism akin to deletion, both of which appear learnable from exposure to alike coordination alone.

Jiamu Luo, Shane Steinert-Threlkeld · 0 citations
Review Open access Aug 2026

Unifying the structures of language in a neural population code

It is concluded that explaining how language can emerge from neural population codes, in both biological and artificial systems, will not be achieved through the incremental refinement of algebraic-symbolic theories but will demand new theoretical paradigms.

Samuel A. Nastase, Zaid Zada, A. Goldberg et al. · 1 citation
Jul 2026

Instruction-Tuned Models Locally Reuse Human Syntax More Than Humans Do

Syntactic convergence (the tendency of speakers to adapt in language towards the grammatical profiles of their interlocutors) is a well-documented feature of human dialogue widely considered to operate below conscious awareness. Whether large language models exhibit analogous syntactic convergence toward human users relative to human baselines and across a broad range of syntactic constructions remains an open question. Using substitution-paradigm data in which model generations replace one speaker's turns in pre-existing human dialogues, this study measures turn-adjacent reuse of context-free grammar (CFG) rules across sixteen open-weight Llama and Gemma models (1B-70B, pretrained and instruction-tuned) at 1,901 matched positions per model. Every model showed greater CFG-rule overlap with the preceding human turn than with a sampled unrelated human prime, and in every model this actual-versus-random difference was larger for lower-frequency rules. Each instruction-tuned model also showed greater natural-output overlap with the actual prime than the human response it replaced, and all eight matched architecture pairs exhibited greater actual-prime overlap after instruction tuning. However, relative to pretrained variants, instruction-tuned outputs overlapped more with unrelated primes, showed a smaller actual-versus-random increment, and had lower conditional rule-reuse odds once target rule-set size was held constant. In exploratory analyses, each model exhibited greater mean lexical and semantic similarity to the preceding turn than the matched human responses did. Instruction-tuned models additionally produced responses with greater mean semantic similarity than their pretrained counterparts in all eight architecture pairs, whereas the lexical similarity results were more heterogeneous.

Zandi Eberstadt · 0 citations
Review Open access 2026

Large Language Models as Distributional Baselines for Language Tasks

The central contributions of this paper articulate the conditions under which distributional predictability threatens the internal validity of an experiment and provide concrete recommendations for how to control for this potential confound.

Sean Trott, James A. Michaelov, Cameron R. Jones et al. · 0 citations
Preprint Aug 2026

Rethinking and formalising the state across languages: a unified computational learning theory account

It is suggested that the state is a systemic, context-dependent morphosyntactic mechanism that selects grammatical templates across synthetic languages and constitutes one instance of a broader class of syntactically conditioned dependencies that also includes agreement and grammatical case.

M. E. Idrissi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.