Skip to content
Open access

Token Trajectories as Knowledge Representation in Large Language Models

Jul 2026 · Philosophies · Vol 11, pp. 116 · 0 citations · 67 references

TL;DR

An interpretation of the appearance of this phenomenon as an emergent effect of a complex system, such as LLMs, is proposed, based on the observation of the computational trajectory of computational units (vectors) which represent discrete semantic units, tokens.

Abstract

The subject of this paper is the problem of knowledge revealed by the emergence of advanced natural language processing technology, called large language models (LLMs). It proposes an interpretation of the appearance of this phenomenon as an emergent effect of a complex system, such as LLMs. The meaning of this interpretation is based on the observation of the computational trajectory of computational units (vectors) which, on the other hand, represent discrete semantic units, tokens. To justify the autonomy and relevance of this interpretation of knowledge, the paper invokes the concept of discursive space, which allows, in particular, for the reliance on language as a container of knowledge and for the generalization of the idea of knowledge to other semantic environments. To describe more general units of knowledge, the strong version of the theory proposes the introduction of the institution of gnosemes, of which discourse is a special case. Gnosemes find application in the case of LLMs. Due to the different ontical contexts of the entities described, the paper proposes a transdisciplinary approach, the axis of which is the field of social sciences.

Read PDF

Similar papers

Open access 2026

Comparison of the conceptual framework node of knowledge (NOK) with large language models (LLM)

Comparisons of large language models with a system built using the Node of Knowledge conceptual framework revealed similarities and differences between the systems, which are presented in this paper.

Martina Ašenbrener Katić, Marina Rauker Koch, A. Jakupović · 0 citations
Open access Jul 2026

L-language and N-language

This article presents an ontological taxonomy of language developed from the perspective of general and applied linguistics. Corresponding to the noncountable form of the word ‘language’ is a family of entities that figure in the realization of the universal, evolved sociocognitive capacity (‘L-language’). Yet the dominant conceptualization of language entails a fundamentally distinct ontological category expressed by the countable form of the word, as a set of monolithic, named languages (‘N-language’). This category is believed to be constituted by sets of determinate social norms, viewed as the inalienable possessions of so-called ‘native’ speakers. It arose relatively recently, as powerful cultural groups ideologized their conceptualization of how they used language as a shared resource, resulting in an ontological shift to language as a characteristic of the nation-state. The fundamental distinction is not adequately acknowledged in much language scholarship, thereby compromising its ontological rigour. To provide some clarity, I outline an approach to language ontology that identifies multiple entities across different domains and allows for distinct linguistic systems to be viewed as collections of lexico-grammatical constructions, independent of peoples and places.

Christopher J. Hall · 0 citations
Jul 2026

Relation Geometry in Semantic Space of Language Models

The results empirically show that relation geometry is not equally well-represented for all relations in semantic space, suggesting that there is a difference in how well semantic relations might be learned from distributional information alone.

Zhihan Cao, Hiroaki Yamada, Simone Teufel et al. · 0 citations
Open access Aug 2026

On the Space of Language and the Language of Spaces in Lexicographic Discourse: Modelling and Interpretation Options

This article presents the results of the study of spatial concepts in the metalanguage of linguistics of the late twentieth and early twenty-first centuries, when the issues of ‘the space of language and the language of spaces’ (E. S. Kubryakova) acquired important theoretical and practical significance and relevance and attracted the attention of prominent linguists remaining in the spotlight to this day. Also, the author examines the dynamics of the formation of the paradigm of terms for spaces of language that are studied theoretically and employed in practical analysis. These terms are actively used as lexicographic parameters in the modelling of multi-dimensional linguistic sets of varying levels and importance within the structure of ideographic dictionaries of the Ural Semantic School. In the spatial dimension, different cognitive strategies for representing the world of experience have been identified; these strategies shape the image of the linguistic worldview. The article describes the constructive aspect of ideographic dictionaries, together with the cognitive strategies governing their internal organisation, demonstrating the role of spatial measurement in that organisation. It is shown that the construction of integrated linguistic spaces is based on the principal modelling rule, i.e. the horizontal-vertical arrangement of sets of units within the external global structure of ideographic dictionaries. This principle is applied in different respects: in the semiological aspect as lexical and semantic linguistic spaces; in the semasiological aspect as denotative and conceptual spaces; and in the system-ideographic aspect as lexico-semantic, semantic-derivational, functional-semantic, denotative-ideographic, conceptual-ideographic, cognitive-discursive spaces, and other analogous sets of component units. These types of spaces interact in lexicographic discourse during the compilation of dictionaries of various kinds, thereby shaping their genre structure. Particular attention is paid to the complex of lexicographic parameters of different ideographic dictionaries and their constituent sets of units, whose composition and content are related to the cognitive mechanisms of spatial measurement of various kinds.

L. Babenko · 0 citations
Aug 2026

The Cybernetic Order-word: Tensors, Tensions and LLM Vector Spaces

Why what is really a matter of data analytics and statistical prediction is so readily assumed to be a display of real intelligence and even emergent cognition is explored by genealogically tracing the relationship between machines, organisms and language.

Chantelle Gray · 0 citations
Open access Jul 2026

Large Language Models as Phenomenotechniques for Linguistics

With the help of Bachelard’s concept of phenomenotechnique – meaning an artificial, technological set-up that forces a certain part of reality to reveal itself in a way that can be scientifically observed – this paper tries to theorize large language models (LLMs) as a kind of “epistemological mirrors” that make symbolic language (as a foundation of culture) observable in an unprecedented way. Consequently, LLMs have (besides their known and attested commercial value) immense epistemological value for the humanities since they allow for an alternative way of studying language learning, generation, and structure that can supplement classic linguistic techniques. LLMs can thus be used to reassess the relevance of existing linguistic theories and update them as well as to start developing new theories of language. The paper develops its thesis through a refutation of a common objection to LLMs – that they have no “true” understanding of language since they possess no relation to the outside world – and by showing that real-world relations are irrelevant for human language as well. Consequently, linguistic theories (such as structuralism) that understood language outside of such a framework are once again relevant for a new understanding of symbolic language and culture that LLMs allow.

Primož Krašovec · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.