Skip to content

Verbalizable Representations Form a Global Workspace in Language Models

Jul 2026 · arXiv.org · Vol abs/2607.15495 · 40 citations · ⚡ 6 influential
Computer Science

TL;DR

Evidence that an analogous functional distinction has emerged in large language models is presented, showing that language models maintain a small, privileged set of representations bearing some of the functional hallmarks of conscious access, and that decoding these representations sheds light on ongoing cognitive processes.

Abstract

Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning. In this paper, we present evidence that an analogous functional distinction has emerged in large language models. Using a new interpretability technique, the Jacobian lens, we identify the representations a model is poised to verbalize at any point in its processing. These representations, which we collectively call the J-space, exhibit the functional properties characteristic of a global workspace: their contents can be reported, deliberately summoned and held, used to carry the intermediate steps of silent reasoning, and passed as arguments to arbitrary downstream computations, while automatic processing such as text parsing and routine inference proceeds without them. The J-space also has structural signatures that global workspace theory associates with conscious access: it carries coherent content only in an intermediate band of layers, holds on the order of tens of concepts at a time, and is broadcast by the model's weights more widely than other representations. These properties make it a practical window into a model's unspoken thinking. In alignment audits, it reveals strategic deliberation, evaluation awareness, and trained-in misaligned dispositions that never appear in the model's outputs. We find that post-training installs the Assistant's point of view in the workspace, and we introduce counterfactual reflection training, which improves behavior by training only what a model would say if interrupted and asked to reflect. These results indicate that language models maintain a small, privileged set of representations bearing some of the functional hallmarks of conscious access, and that decoding these representations sheds light on ongoing cognitive processes.

View source

Similar papers

#natural language process... Preprint Sep 2026

Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs

Reasoning in large language models unfolds through diverse functional operations, such as problem formulation, goal decomposition, and deduction. Although these operations are explicitly distinguished in text, little is known about how they are geometrically organized in representation spaces. To this end, we investigate whether distinct reasoning operations exhibit corresponding geometric structure in hidden representations. We find that operations are separable in held-out representations, with separability peaking in middle layers, and verify that this structure is not explained by lexical or positional confounds. Across layers, token-wise operation-alignment becomes more distributed over spans, while identical surface tokens are represented differently depending on the operation of its surrounding chunk. Attention-masking interventions further show that operation-aligned representations at chunk onset depend on preceding reasoning context. Consequently, our work demonstrates that language models maintain representational correspondence between linguistic reasoning expressions and their internal geometric structures. Code and project materials are available at https://github.com/naver-ai/beneath-cot.

Seogyeong Jeong, Jaehui Hwang, Dongyoon Han et al. · 0 citations
Preprint Aug 2026

Conscious Access as Continuous-to-Discrete Translation

The scientific study of consciousness frequently stalls on ontological debates regarding the"Hard Problem."This paper proposes a pragmatic pivot. Rather than asking what consciousness is metaphysically, we ask how modeling conscious access as a specific computational transformation may address existing bottlenecks in neuroscience and artificial intelligence. We introduce the Continuous/Discrete (C/D) framework, which holds that the brain implements two distinct processing regimes: System C, a distributed sensory-motor network operating over continuous, high-dimensional manifolds, and System D, a centralized engine structured around discrete, scale-invariant symbols. We argue that conscious access requires a structure-preserving translation between these regimes, which maps localized continuous states onto discrete symbolic tokens, coupled with an inverse projection that grounds those tokens back into sensorimotor dynamics. By formalizing conscious access as this continuous-to-discrete conversion, we derive a unified set of testable predictions centered on representational geometry, specifically, on a measurable collapse from graded similarity structures to low-dimensional categorical equivalence classes. These predictions explicitly differentiate our account from Global Neuronal Workspace Theory, Integrated Information Theory, Predictive Processing, and Higher-Order Theories, shifting the focus from ontological status to computational mechanism. Beyond neuroscience, the framework provides a principled architecture for neuro-symbolic artificial intelligence. We argue that treating conscious access as translational computation offers a pragmatic, empirically tractable pathway forward, clarifying what conscious states functionally accomplish without requiring resolution of the hard problem of phenomenology.

Tianming Yang · 0 citations
Review Open access Aug 2026

Unifying the structures of language in a neural population code

It is concluded that explaining how language can emerge from neural population codes, in both biological and artificial systems, will not be achieved through the incremental refinement of algebraic-symbolic theories but will demand new theoretical paradigms.

Samuel A. Nastase, Zaid Zada, A. Goldberg et al. · 1 citation
Jul 2026

J-CoT: Chain-of-Thought in J-Space

Under matched backbone and inference settings, J-CoT-Zero matches or exceeds the strongest evaluated latent-reasoning baseline on every benchmark, while J-CoT-Train obtains the highest score across the evaluated mathematical, scientific, coding, and structured path-reasoning tasks.

Jun-De Wu, Jiayuan Zhu, Feng-Lin Liu et al. · 1 citation
Open access Jul 2026

Linguistic structure and probability are jointly encoded in high gamma power

During speech comprehension, the brain dynamically infers a hierarchy of increasingly abstract representations from the sensory input. An important step in the inferential hierarchy is the combination of words to form phrases and sentences. Whether this process is driven primarily by statistical patterns in the linguistic input, or by a mechanism that combines words into hierarchical representations, is a subject of considerable debate that has regained importance with the arrival of large language models. This study investigates whether local cortical activity (high gamma power; 70-150 Hz) from intracranial recordings is jointly modulated by lexical probability and syntactic structure; and whether lexical probability affects the inference of syntactic structure. To this end, an open dataset of electrocorticography recordings is analyzed with multivariate temporal response functions and a model comparison approach. The results indicate that high gamma power is sensitive to multi-word estimates of constituency structure and lexical probability estimates, both in isolation and jointly. The temporal response functions suggest that syntactic structure building depends on interregional communication between regions connected through dorsal- and ventral streams. Furthermore, the study provides evidence that bottom-up syntactic information is less likely to be encoded by neural populations that strongly code for lexical probability measures, while top-down syntactic information shares neural resources with lexical uncertainty. We suggest that lexical uncertainty modulates the weighting of anticipatory structural information. With this, the current study supports models that suggest that cues are leveraged flexibly in a feed-forward and feed-back fashion during speech comprehension.

S. Slaats, A. Hervais-Adelman · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.