Back to feed
Preprint

Constructing Parallel Multidimensional Chromatic Lexicons for Corpus-Assisted Analysis of Russian and English Texts

Aug 2026 · 0 citations · 31 references
Computer Science

TL;DR

The development of two multidimensional chromatic lexicons for Russian and English is described, with the main contribution of this study being a transparent and reusable procedure for constructing and applying multilingual chromatic lexicons.

Abstract

This article addresses the relative scarcity of research tools for the corpus-assisted linguistic analysis of colour terms in literary texts. It describes the development of two multidimensional chromatic lexicons: one for Russian (224 entries) and one for English (141 entries). Lexicon construction involved sourcing colour vocabulary from specialised resources and research literature, comparing the two language inventories, manually checking translated candidates, and addressing language-specific morphological features. In addition to identifying colour terms and visual descriptors, the lexicons classify entries according to hue, saturation, and temperature. To demonstrate their practical application, a pilot study was conducted on purposively sampled corpora of poetry by Andrei Bely (20,373 tokens) and Emily Dickinson (28,479 tokens). All retrieved matches were checked in context and classified as Confirmed_chromatic, Ambiguous_visual, or Excluded. The analysis was implemented in two main stages: a strict analysis including confirmed chromatic lexis only, followed by a sensitivity analysis incorporating both confirmed and ambiguous chromatic lexis to determine whether coding decisions about borderline cases affected the main findings. The quantitative results indicated marked differences in the use of colour terms, visual descriptors, hue, saturation, and temperature. Specifically, the analysis revealed that confirmed chromatic terms occurred 3.4 times more frequently in the sampled Bely corpus than in the Dickinson corpus. These findings demonstrate the analytical value of a multidimensional approach, with the main contribution of this study being a transparent and reusable procedure for constructing and applying multilingual chromatic lexicons.

View source

Similar papers

Open access Jun 2026

Lexical competition in Kazakh–Russian bilingual media: A corpus-based quantitative study

In recent years, corpus linguistics has provided new opportunities for studying language contact using quantitative and data-driven techniques. For the Kazakh language, situated in a multilingual setting involving Kazakh, Russian, and English, lexical interference represents an important sociolinguistic issue. This study investigates lexical competition from quantitative and functional standpoints, with particular attention to frequency, distribution, and variation in Kazakh-language mass media under Kazakh-Russian language contact. Using corpus methods, the study shows that large collections of texts make it possible to identify borrowed lexemes and competing equivalent forms. Borrowed items may coexist with native equivalents, producing measurable differences in frequency and stylistic preference. In some instances, foreign lexemes are more frequent than standardised national alternatives in the analysed corpus, although established Kazakh norms may continue to coexist with such borrowings. To address this issue, the study introduces the Lexical Competition Coefficient (LCC), a descriptive numerical indicator that measures the relative frequency of competing lexical equivalents in corpus data. When the Kazakh item predominates, the LCC exceeds one (k > 1), indicating that the Kazakh variant is more frequent in the corpus. By contrast, a coefficient below one (k < 1) indicates that the borrowed variant is more frequent. The findings indicate that lexical competition in Kazakh-language media is uneven across lexical pairs, with some borrowed forms substantially outnumbering Kazakh equivalents while other Kazakh variants remain dominant in corpus usage. The study combines scholarship on language contact and quantitative corpus analysis and proposes a systematic procedure for assessing lexical competition in corpus data.

A. Ormanova, M. L. Anafinova, Kaziza Kozhagulova · 0 citations
Open access Jul 2026

Cognitive Misalignment and Categorical Reconstruction of Qi in English Translation

The translation of culture-specific items is difficult not merely because of lexical gaps, but because languages often organize experience through different underlying cognitive categories. This study examines the Chinese core concept qi (Gas), a radial category rooted in embodied experience and centered on the prototype of vital energy, and explains why it is frequently fragmented in English translation. A bilingual database of 49 high-frequency qi-compounds was built from the Chinese Proficiency Grading Standards for International Chinese Language Education and checked against the Modern Chinese Dictionary and Oxford Learner's Dictionaries. The static morphological comparison was then supplemented by qualitative analysis of authentic translations from medical classics, philosophy, literary theory, ethical discourse, and fiction. The results show a systematic mismatch: Chinese organizes the qi-word family through a shared morpheme and modifier-head construction, with 34 of 49 items placing qi in final position and 37 of 49 exhibiting modifier-head structure. By contrast, English renders these items through morphologically unrelated simplexes, derivatives, phrases, and one compound, thereby distributing a continuous Chinese category across discrete lexical domains such as air, strength, anger, courage, integrity, and atmosphere. Translation examples further reveal metaphorical rupture when the force, container, and flow schemas of qi are replaced by static English entities. To address this problem, the study proposes a cognitive compensation framework consisting of category transplantation, semantic compensation, category re-implantation, and context-sensitive modulation. The framework contributes to cognitive translation studies and offers practical implications for translating and teaching culturally dense Chinese key concepts in applied language-learning contexts.

Ting Zhang, Zihan Yu · 0 citations
Open access Jul 2026

Morphologically Annotated Lexical Decision Data for 12,242 Czech Word Forms

We present CzeLeD, a large-scale lexical decision dataset for Czech containing 12,242 lemmas and inflected word forms derived from the HeCz self-paced reading corpus and supplemented with pseudowords. A total of 1,977 participants contributed over 849,000 trials, accompanied by demographic data and extensive linguistic annotation, including lemma and word-form frequencies, neighborhood density, phonotactic and orthotactic probability. Alongside the trial-level dataset, we provide lexical decision norms summarizing reaction times and accuracy using a transparent and reproducible trimming procedure. Data validation drew on both response accuracy and reaction times. Participants showed high overall accuracy (94.79%), and reaction times exhibited the expected strong inverse relationship with word form frequency. Importantly, CzeLeD can be directly linked to the HeCz corpus, enabling systematic comparison between word recognition in isolation and in sentential context. By combining broad lexical coverage, rich annotation, and seamless integration with an existing context-bound processing corpus, CzeLeD offers a powerful, reusable resource for investigating lexical processing, morphological complexity, and contextual effects in Czech and beyond.

J. Chromý, Markéta Ceháková, Mikuláš Preininger et al. · 0 citations
Review Open access Jul 2026

Natural language processing for Arabic poetry analysis and generation: a systematic review

Poetry is a unique form of expression valued for its role in preserving cultural heritage. Analyzing Arabic poetry is time-consuming and requires a high level of linguistic expertise; therefore, computational methods are useful, as they enable large-scale, extensive, and systematic analysis of poetry, thereby improving its accessibility for researchers and students. This article presents the first systematic review of natural language processing (NLP) and machine learning (ML) approaches for Arabic poetry. It addresses the question of which research tasks, methodologies, datasets, and evaluation approaches have been applied to Arabic poetry, and which trends and research gaps can be identified in the existing literature. In accordance with Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA), the author conducted an exhaustive search across six major academic databases (ACL Anthology, IEEE Xplore, ACM, SpringerLink, Science Direct, and Google Scholar) for relevant studies published between January 2010 and May 2025. Eligibility was evaluated in several phases, and re-examination was conducted to ensure accuracy. The author performed task-level categorization, extracted key characteristics from each study, synthesized the findings, and presented them in tables and figures to highlight the main trends and research gaps in the literature. This study presents the first structured task-level synthesis of the field, identifying methodological trends, detecting evaluation inconsistencies, and highlighting research gaps that have not been critically consolidated before. Furthermore, the author assembled a comprehensive collection of available datasets and resources to promote standardized assessment.

Haifa Alharthi · 0 citations
Open access Aug 2026

Exploring Mixed Copies: Evidence From English-Estonian Bilingual Speech

Recently, contact linguistics has become increasingly interested in multiword units. At the same time, the code-copying framework (CCF) includes the notion of mixed copies (MCs) that are in-between global copies (‘borrowing’) and selective copies (‘structural change’) and illustrate the transition between the lexicon and grammar. The research question is: What types of MCs occur in English-Estonian bilingual speech? The data were transcribed, and MCs identified, annotated, and classified according to their structure. English items were searched for in Estonian dictionaries to establish their Estonian equivalents or conventionalization of such items. The frequencies of MCs and their Estonian equivalents were also searched on Google to determine whether the MCs occur outside the corpus. Three datasets were analysed: written texts from 44 blogs (385,124 tokens), spoken data from 10 vlogs (117,555 tokens), and 8 podcasts (77,277 tokens). Quantitative analyses of the various MC types were conducted, followed by a qualitative analysis of representative examples. Compound nouns constitute the majority of MCs, followed by idioms, phrasal compounds, and a small number of compound verbs. No frame-changing MCs (i.e., MCs resulting in grammatical change) were attested. Since compound nouns and analytic verbs occur in both languages, structural similarity may be a facilitating factor in copying. The notion of MCs is not widely used. Research typically focuses on particular types of items (e.g., compound nouns or verbs); here, however, the question is reversed: which types of items yield MCs? It was established that the proportion of MCs in the data is comparable to that of selective copies. Within MCs, the globally copied element renders the remaining part more specific, highlighting the importance of meaning in contact-induced language change. MCs are also present on the Estonian internet and, in some cases, outnumber their Estonian equivalents, if such equivalents exist.

A. Verschik, H. Kask · 0 citations
Open access Jul 2026

Evaluating the limits of machine translation for poetry: a multidimensional framework

The results show that LLMs and Google Translate consistently outperform specialized MT systems in terms of fluency, meaning preservation, and lexical-thematic alignment.

Beatriz Ribeiro Borges, P. H. R. Gabriel, E. Faria · 0 citations