Aug 2026· Bulletin of the National Technical University KhPI Series Actual problems of Ukrainian society development· pp. 59-65· 0 citations
TL;DR
It is argued that automated detection should not replace expert linguistic analysis but should serve as an auxiliary tool for evaluating the probable origin of a text in media linguistics, fact-checking, and educational practic.
Abstract
The article examines linguistic features of Ukrainian socio-political news texts generated by large language models and methods for their automated identification. The aim is to substantiate lexical, stylistic, compositional and semantic markers that may indicate AI-generated text and to outline a computational-linguistic detection framework. The study proposes combining philological interpretation with NLP procedures: preprocessing, lexical diversity assessment, clustering, vectorization and transformer-based classification. It is argued that automated detection should not replace expert linguistic analysis but should serve as an auxiliary tool for evaluating the probable origin of a text in media linguistics, fact-checking, and educational practic
It was concluded that texts generated by artificial intelligence constitute a separate linguistic phenomenon with its own set of characteristics, which requires a special typology and a flexible, updatable analysis methodology.
L. Kravets, Viktória Stefuca, N. Libak et al.· Computer and Decision Making...· 0 citations
The results demonstrate that linear discriminative models and appropriate lexical feature engineering can provide a very accurate and interpretable baseline for the development of natural language processing algorithms for Arabic.
Hamood Mohammed Alrumhi, Muhammad Asshad, Amjed Abbas Ahmed et al.· JOIV: International Journal...· 0 citations
Analysis of AIGC texts points out that the complexity of AI text primarily stems from its mechanism of selecting vocabulary based on probability distributions, which favors longer words, abstract nouns, and words with high semantic content, thereby forming a highly compact linguistic surface.
Lulu Chen· Lecture Notes on Language an...· 0 citations
Semantic analysis has become a central challenge in natural language processing, driven by exponential growth in digitized textual data and the need for automated content processing across multiple applications including machine translation, text classification, sentiment analysis, and information retrieval. However, while semantic analysis methods are well-developed for resource-rich languages such as English, morphologically complex languages like Uzbek suffer from deficiencies in annotated corpora, lexical-semantic resources, and high-quality vector models – a gap amplified by governmental initiatives in digital economy development and national language technology advancement. This section grounds semantic analysis in the distributional semantics hypothesis principle that words exhibiting similar contexts possess similar meanings – thereby recasting the problem as a geometric challenge within continuous vector spaces. Two principal mathematical strategies are formalized: (1) prediction-based models (word2vec: CBOW/Skip-gram), which optimize context prediction objectives, and (2) count-based models (GloVe), which leverage global co-occurrence statistics through matrix factorization. Both project high-dimensional word co-occurrence relationships into low-dimensional dense vector spaces, enabling semantic analogy representation. For resource-scarce languages like Uzbek, cross-lingual embedding alignment (Procrustes optimization) enables semantic knowledge transfer from resource-rich languages, facilitating shared semantic spaces across the Turkic language family. The section concludes with formal problem specification: given vocabulary V and corpus C, semantic analysis is formalized as (1) a mapping problem preserving distributional properties, (2) an optimization problem minimizing loss through gradient-based methods, and (3) an evaluation problem assessing quality through semantic similarity, analogy, and downstream NLP task performance.
D. Akhmedjanova· Международный Журнал Теорети...· 0 citations
A statistically significant increase in lexical density and frequency of cohesive markers has been revealed, indicating an increase in information compression and explicit discursive organization of texts, consistent with characteristics described in AI-assisted writing studies.
T. Nedashkivska, I. Varvaruk, M. Podoliak et al.· Journal of Intelligent Decis...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.