Jul 2026· Natural Language Processing· Vol 32, pp. 391-429· 0 citations· 237 references
TL;DR
By categorizing LSC models into generations defined by key components, this investigation suggests that performance breakthroughs have been largely driven by advances in semantic representations, transitioning from count-based models to recent transformer-based approaches.
Abstract
Lexical semantic change (LSC) detection investigates changes in word meaning over time, focusing on language use at the lexical-semantic level from a diachronic perspective. The field has made significant progress over the past two decades, driven by the increased availability of multilingual benchmarks, notable performance improvements, and growing interdisciplinary applications. In this paper, we review the evolution of LSC models and benchmark constructions within the context of popular shared tasks. By categorizing LSC models into generations defined by key components, our investigation suggests that performance breakthroughs have been largely driven by advances in
semantic representations
, transitioning from count-based models to recent transformer-based approaches. Notably, transformer models have established themselves as state-of-the-art by integrating
Word-in-Context
tasks, which emphasize
semantic proximity in context
. Furthermore, we review substantial studies that primarily leverage diachronic word embeddings to explore political, social, and cultural contexts beyond the linguistic domain. Our work provides valuable insights for future model development and encourages further interdisciplinary exploration within digital humanities and social sciences.
Lexical semantic change (LSC) is commonly modelled through vector-space representations, but these approaches often provide limited insight into which aspects of usage are changing. Diachronic corpus research instead examines interpretable dimensions such as syntactic behaviour, morphology, and constructional patterns, but typically through separate analytical workflows. We present SynFlow, an open-source toolkit for multidimensional diachronic analysis of linguistic usage. SynFlow converts linguistic observations into period-specific distributions and applies a shared workflow across dependency-based co-occurrences, morphological features, constructional configurations, and externally derived representations such as Frame Semantics. It supports different distance measures, together with value-level decomposition, statistical testing, and incremental clustering of lexical fillers. We demonstrate SynFlow through a qualitative case study of the German adjective viral, showing how a single semantic development is reflected across syntactic, lexical, constructional, and morphological dimensions. We further report previously published results on SemEval-2020 Task 1 to situate the performance of these representations relative to existing lexical semantic change detection systems.
Bach Phan-Tat, K. Heylen, Dirk Geeraerts et al.· 0 citations
Dynamic topic models capture evolving word distributions, but traditional coherence metrics may fail when vocabulary changes while semantic meaning persists. We evaluate 120 topics from CoNTM and DLDA across NYT, DBLP, and arXiv, using three human annotators and Low, Medium, and High lexical-change categories. Traditional temporal coherence shows highly variable agreement with human judgments ($\rho$=-0.256 to 0.614). In contrast, LLM-based semantic similarity agrees strongly with human semantic judgments for CoNTM on NYT ($\rho$=0.609), DBLP ($\rho$=0.721), and arXiv ($\rho$=0.502), but is less consistent for DLDA. Lexical-change stratification reveals variation hidden by aggregate evaluation. We therefore advocate lexical-change-aware evaluation, jointly reporting traditional coherence and LLM-based semantic measures as complementary rather than interchangeable signals.
Semantic analysis has become a central challenge in natural language processing, driven by exponential growth in digitized textual data and the need for automated content processing across multiple applications including machine translation, text classification, sentiment analysis, and information retrieval. However, while semantic analysis methods are well-developed for resource-rich languages such as English, morphologically complex languages like Uzbek suffer from deficiencies in annotated corpora, lexical-semantic resources, and high-quality vector models – a gap amplified by governmental initiatives in digital economy development and national language technology advancement. This section grounds semantic analysis in the distributional semantics hypothesis principle that words exhibiting similar contexts possess similar meanings – thereby recasting the problem as a geometric challenge within continuous vector spaces. Two principal mathematical strategies are formalized: (1) prediction-based models (word2vec: CBOW/Skip-gram), which optimize context prediction objectives, and (2) count-based models (GloVe), which leverage global co-occurrence statistics through matrix factorization. Both project high-dimensional word co-occurrence relationships into low-dimensional dense vector spaces, enabling semantic analogy representation. For resource-scarce languages like Uzbek, cross-lingual embedding alignment (Procrustes optimization) enables semantic knowledge transfer from resource-rich languages, facilitating shared semantic spaces across the Turkic language family. The section concludes with formal problem specification: given vocabulary V and corpus C, semantic analysis is formalized as (1) a mapping problem preserving distributional properties, (2) an optimization problem minimizing loss through gradient-based methods, and (3) an evaluation problem assessing quality through semantic similarity, analogy, and downstream NLP task performance.
D. Akhmedjanova· Международный Журнал Теорети...· 0 citations
Recent developments in NLP and web-scale document analysis have increasingly emphasized the importance of interpretability and contextual dependence in semantic representations. Although modern word embeddings achieve remarkable empirical performance, their semantic structure is often difficult to interpret, since meaning is encoded through latent geometric relations in high-dimensional spaces. This paper discusses an alternative conceptual framework based on explicit contextual semantic relations. Building on ideas from distributional semantics, co-occurrence analysis, and fuzzy set theory, the study revisits semantic projections and related count-based representations as interpretable directional semantic structures for semantic analysis in document corpora and web-based information environments. In this setting, several classical association measures, including PMI and related transformations, may be understood as derived from simpler conditional semantic projections. The methodology is illustrated through a comparative analysis of semantic associations related to “ChatGPT” across general web-scale data and specialized scientific repositories. Our results demonstrate that semantic projections effectively capture persistent contextual structures while remaining sensitive to corpus-specific discourse communities. The resulting perspective emphasizes interpretability, asymmetry, contextual dependence, and direct empirical meaning as central principles for semantic representation.
Mabel López-Bordao, Antonia Ferrer-Sapena, Pablo Lara-Navarra et al.· Information· 0 citations
A methodical investigation of temporal prompting techniques for LLM-based EL is presented, and it is demonstrated that explicit temporal prompting can reduce drift mistakes by up to 40% using a dataset of temporally-sensitive mentions linked with Wikidata snapshots.
A data-centric analysis of semantic knowledge acquisition in word embeddings, focusing on word analogy and semantic similarity shows that, for relational semantics, training-data quality outweighs quantity, and that simple proxy models remain a practical, interpretable tool for efficient data selection.
Aishwarya Jadhav, Mark Anderson, José Camacho-Collados et al.· Neural computing & applicati...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.