The iterated learning model was introduced to investigate language evolution: the way in which the characteristic properties of human languages have been shaped, at least partly, by repeated transmission from one language user to another. The key finding is that language compositionality can arise spontaneously as a consequence of language being passed repeatedly through a language learning bottleneck. Here we explore how changing the frequency of different meanings, so that some meanings occur much more frequently than others, affects the character of its compositionality. We find that, as observed in natural languages, high-frequency meanings can escape the pressure to conform to the grammar that characterizes lower-frequency meanings. However, when the frequency structure is instead imposed on parts rather than on whole meaning vectors, the language fails to transmit across generations. This occurs despite the fact that the most frequent elements are reliably learned. These results suggest that frequency can shape emergent linguistic structure only when the frequency distribution is defined over form-meaning units that learners can acquire holistically. When frequency is instead distributed over smaller units, it fails to support the relational structure required for compositional generalisation, thereby preventing stable language transmission.
All languages have statistically coherent subsequences (e.g., words) whose frequency distribution follows a power law. These properties facilitate language learning in humans, making them good candidates for arising through cultural transmission as a way to help faithful transmission across generations. Recently, both properties were found in whale song, which is also culturally transmitted, leading to the strong prediction that they should be found wherever complex sequential signaling is culturally transmitted. Here, we use the same tools used to analyze human data and whale song to reveal that culturally transmitted Bengalese finch song also has statistically coherent subsequences whose distribution follows a power law. We additionally show that statistical coherence increases over development, but the power law is present throughout, suggesting that it reflects a fundamental principle of learned representations. Finding these parallels between evolutionarily distant species illustrates the importance of cultural transmission in shaping communication and suggests that core properties of language arise through convergent evolution.
Simon Kirby, Kazuo Okanoya, Ellen C. Garland et al.· Science Advances· 1 citation
We build rule-based emergent language (EL) agents using form–meaning mappings induced from ELs (“morphological phrasebooks”) and test their communicative performance in the EL environment with its neural network agents. This contributes three things: First, it assesses the quality of the morphemes discovered by the induction algorithm in situ , which we find to be effective for communicating in the EL. Second, it allows us to uncover morphosyntactic properties of EL through ablating the algorithms which induce and utilize morphemes, showing that the ELs rely on repetition as well as morpheme ordering to convey meaning. Third, we find that the normalized pointwise mutual information of forms and meanings in the mor-phemes serves as a metric of compositionality that is more closely correlated with the ability of the phrasebook-agents to “speak” and “hear” an EL than existing metrics such as topographic similarity.
Brendon Boldt, David R. Mortensen· Annual Meeting of the Associ...· 0 citations
Large language models (LLMs) have mastered human language in ways that no previous computational system has. While rule-based, symbolic systems sufficed for constrained, well-defined problems, they were not able to accommodate the context-sensitive expressivity of natural language. LLMs instead use statistical learning to encode the diversity of linguistic structures into a unified high-dimensional embedding space. Strikingly, this context-driven, distributed representation closely parallels neural population codes, suggesting that the human language system may have converged on a similar computational strategy. Drawing on a growing body of work at the intersection of artificial intelligence and cognitive neuroscience, we show that LLMs can serve as cognitively plausible models of the neural computations supporting language in the human brain. We conclude that explaining how language can emerge from neural population codes, in both biological and artificial systems, will not be achieved through the incremental refinement of algebraic-symbolic theories but will demand new theoretical paradigms.
Samuel A. Nastase, Zaid Zada, A. Goldberg et al.· Neuron· 0 citations
Human language exhibits lawful structure at the level of words (frequency, vocabulary growth) and word pairs (co-occurrence across distance). Here we show that the arrangement of words in sequence -- a central determinant of meaning -- obeys a comparable law. Using large language models as probabilistic probes, we measured the reduction in target perplexity conferred by prior context at distance d beyond that of the same words scrambled; this difference, the contextual persistence function P(d), isolates the influence of arrangement. Across ten corpora spanning six language families and written and spoken modalities, P(d) decayed approximately as 1/d ($P(d) \propto d^{-\alpha}$, mean $\alpha = 1.04$; median $r^2 = 0.96$). The effect vanished in scrambled and synthetic controls, replicated across independent probes, and did not appear in genomic or protein sequences under domain-native models. An exponent near 1 distributes contextual influence approximately uniformly across logarithmic timescales. The results establish a scaling law of contextual persistence in human language.
A scale-dependent transition between two ID regimes is found: at low lexical diversity, conditions with fewer unique final words produce higher ID, while at high lexical diversity, this ordering reverses, and conditions with more unique words produce higher ID.
Arwa Osman, Marco Baroni, Iuri Macocco· 0 citations
What happens when large language models (LLMs) begin to ‘speak’ like humans? Is it an instance of robot intelligence or reason? Or is it something else entirely? In this article I explore why what is really a matter of data analytics and statistical prediction is so readily assumed to be a display of real intelligence and even emergent cognition by genealogically tracing the relationship between machines, organisms and language. It becomes clear that machine–organism metaphors have a long history that can be traced back to René Descartes, which once again became prominent during the cognitive revolution and the subsequent development of cybernetics. Of interest is the role of Chomskyan linguistics in this history and critiques thereof by cognitive linguistics who argue for embodiment but retain some functionalist views, such as the primacy of mental representations. In recent work on enaction this is discarded and, as I show, enactive work on linguistic bodies brings us close to a Deleuzo-Guattarian understanding of language as assemblages of enunciation. Enactivists do not, however, have an explicit theory of technicity, which is where Bernard Stiegler, Gilles Deleuze and Félix Guattari provide correctives. Taken together, linguistic bodies, alongside Stiegler's theorisation of grammatisation and mnemotechnics, and Deleuze and Guattari's understanding of the translatability of language and of order-words as tensors provides us with an especially sophisticated framework for thinking about the Bayesian ordering of the world.
Chantelle Gray· Deleuze and Guattari Studies· 0 citations