Skip to content

Comparing Semantic Navigation in Humans and Large Language Models using Natural Language Processing

Jul 2026 · arXiv.org · Vol abs/2607.12195 · 0 citations · 52 references
Computer Science

TL;DR

It is suggested that human semantic search implements a distinctive balance between local exploitation and global exploration that current model architectures fail to reproduce.

Abstract

Semantic memory retrieval can be conceptualized as navigation through conceptual space. We compared semantic search dynamics between humans and three large language models (GPT-4o, Gemini-2.5-Pro, Claude-Sonnet-4.5) using verbal fluency data. By applying trajectory-based NLP metrics to the items generated by 82 human participants and LLM output across eight temperature settings, we quantified three complementary dimensions: entropy (step size predictability), distance to next (successive semantic steps), and distance to centroid (global dispersion). Humans exhibited higher entropy, larger semantic steps and broader dispersion than all LLMs, indicating more variable and exploratory search. Temperature tuning produced only partial alignments, as individual metrics matched between humans and LLMs at specific settings, but no configuration reproduced the complete human profile (in all dimensions). These findings suggest that human semantic search implements a distinctive balance between local exploitation and global exploration that current model architectures fail to reproduce.

View source

Similar papers

Conference Jul 2026

Investigating Cross-Modal Semantics in Large Language Models Using Concept-Color Associations

Cross-modal correspondences, exemplified by systematic mappings between abstract concepts and colors, illuminate the mechanisms through which humans integrate sensory inputs with conceptual frameworks. Existing theories propose that these associations arise from semantic mediation via metaphors, statistical regularities in language and environment, or overlapping affective or structural features. Trained solely on textual data, large language models provide a unique perspective for examining whether such mappings can emerge from linguistic patterns alone. In this study, we collected color associations for 85 abstract concepts across temporal, alphanumeric, directional, and spatial categories from 260 Chinese university students. These human responses were then compared with outputs from 10 LLM variants across three families, GPT, Deepseek, Doubao, generating 300 responses per concept per model to capture distributional tendencies. Humans exhibited consistent conceptcolor links for 72 of the 85 concepts, with semantically proximate items, such as consecutive seasons or months revealing organized similarities in their color profiles. LLMs, in contrast, produced sharper and more peaked distributions, aligning well with humans on established conventions such as spring-green or A-red, yet diverging on others, for instance associating the concept 0 predominantly with black rather than the human preference for white. Alignment metrics were moderate overall, with the highest levels observed in the GPT family, suggesting that expansive training corpora better approximate nuanced human variability.

Yan Zhang, Hui-Jing Lin, Qi Zhang et al. · 0 citations
Preprint Aug 2026

A Dynamic-Semantics Framework for Grounding Human Referring Expressions in Visual Perceptual Data

Humans converge on shared names for novel, hard-to-describe objects through repeated interaction, a process psycholinguists call lexical entrainment. Leading vision-language models fail at this: recent empirical work documents that they do not shorten references, reuse successful expressions, or maintain stable pact state across turns. We present a framework that addresses the gap by externalizing pact state into three explicit, inspectable sets of referent-object bindings ($\Gamma, \Xi, \Omega$), updated by a dynamic-semantics context-change rule. The symbolic layer sits on top of a lightweight perceptual-alignment pipeline that grounds noisy human referring expressions in crowd-sourced imagery via SIFT homographies and the Universal Quality Index. Evaluated on the Stanford Repeated Reference Game corpus (over 15{,}000 director-matcher utterances on abstract tangram stimuli), the framework places the correct target in its top-5 hypothesis set 83.56% of the time from a single director utterance. Human matcher top-1 accuracy on the same corpus is approximately 77-80%. We also report results on a held-out condition in which obvious tangram-adjacent images are excluded from the retrieved set, which provides a more conservative measurement of the grounding signal. Ablations isolate the contribution of each component: SIFT alignment, UQI, query preprocessing, and image augmentation. The central contribution is the combination: a transparent, auditable symbolic layer that recovers the structure of lexical entrainment turn by turn, paired with a perceptual channel whose behavior can be examined ablation by ablation. We also discuss in detail what the framework does not do. It is not interactive, it does not close the loop with the director, and its retrieval-driven perceptual channel is vulnerable to a class of leakage effects that we quantify and bound rather than wave away.

Joe M. Bingham · 0 citations
Review Jul 2026

Towards High-Level Semantic Intelligence

Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, the development of AI reveals a clear trajectory from simple to complex semantic processing. While early AI systems mainly addressed tasks involving direct and literal semantic perception or expression, contemporary systems are increasingly expected to perform more sophisticated cognitive reasoning, enabling the understanding and generation of High-Level Semantics (HLS). A similar trajectory can also be observed in human cognitive development. We define this transition as the shift from Basic-Level Semantic Intelligence (BLSI) to High-Level Semantic Intelligence (HLSI). However, this issue has not yet been systematically and comprehensively examined in prior work. Motivated by this gap, this survey reviews the development of AI semantic intelligence from the perspective of semantic complexity. We systematically survey existing research on HLS tasks, including humor, sarcasm, metaphor, empathy, persuasion, narrative, and other general HLS phenomena, across text, speech, vision, and multimodal scenarios. Specifically, we summarize data construction methods, modeling and optimization strategies, and evaluation methodologies for both understanding and generation. HLS is essential for advancing AI toward genuinely human-like intelligence. By synthesizing existing methods and insights from the perspective of semantic intelligence, this survey aims to support the continued development of AI toward HLSI.

Xiujie Song, Ge-Fei Yang, Yi-Ning You et al. · 0 citations
Preprint Aug 2026

Divergent large language model predictions from convergent representations in ambiguous word pairs

This work investigates how decoder-only transformers resolve lexical ambiguity through layer-by-layer analysis of three models spanning three parameter sizes, finding that representations become maximally distinct in middle layers, then partially reconverge in late layers, while the KL divergence between their next-token predictions reaches its maximum in the final layers.

K. Scott, Narun Pat, Veronica Liesaputra · 1 citation
#natural language process... Preprint Sep 2026

From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora

We approach a cognitive simulation perspective on episodic and semantic memory in multiple-choice question answering by incorporating text from individual text corpora (ITC) into retrieval-augmented generation and DoRA fine-tuning. We web-crawl the search histories of 515 participants who answered 36 multiple-choice knowledge items and analyze a stratified subsample of 150 participants. For each participant, one DoRA adapter consolidates their ITC into a small language model (SLM) whose baseline correctness falls below the participants'lowest quartile. The adapter measurably writes the ITC into the weights: it fits its own participant's held-out text better than other participants'texts (dz =1.27), an individuality effect that increases with ITC size in rank order. On the generalized knowledge test, however, the adapter adds knowledge rather than alignment with the individual: log-loss match improves, whereas match accuracy under a bias-corrected PMI readout does not, and retrieval adds nothing on top. Our results demonstrate that ITCs can be consolidated into the weights of SLMs, an encouraging basis for individualized tutoring agents, and we discuss how to move from there toward a realistic simulation of episodic and semantic memory at the individual level.

Christoph Wigbels, Ali Abusaleh, M. T. Jansen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.