Skip to content
Open access

Morphologically Annotated Lexical Decision Data for 12,242 Czech Word Forms

Jul 2026 · Scientific Data · 0 citations

Abstract

We present CzeLeD, a large-scale lexical decision dataset for Czech containing 12,242 lemmas and inflected word forms derived from the HeCz self-paced reading corpus and supplemented with pseudowords. A total of 1,977 participants contributed over 849,000 trials, accompanied by demographic data and extensive linguistic annotation, including lemma and word-form frequencies, neighborhood density, phonotactic and orthotactic probability. Alongside the trial-level dataset, we provide lexical decision norms summarizing reaction times and accuracy using a transparent and reproducible trimming procedure. Data validation drew on both response accuracy and reaction times. Participants showed high overall accuracy (94.79%), and reaction times exhibited the expected strong inverse relationship with word form frequency. Importantly, CzeLeD can be directly linked to the HeCz corpus, enabling systematic comparison between word recognition in isolation and in sentential context. By combining broad lexical coverage, rich annotation, and seamless integration with an existing context-bound processing corpus, CzeLeD offers a powerful, reusable resource for investigating lexical processing, morphological complexity, and contextual effects in Czech and beyond.

Read PDF

Similar papers

Open access Jul 2026

How novel are low-frequency words?

This study explores the relationship between low-frequency lexical forms and lexical innovation by examining infrequent and non-lexicalised adjectives formed with the suffix -able (e.g., jokeable, trickable). Employing a corpusbased analysis of the 20-billion-word News on the Web (NOW) corpus (2010–2025), we identified over 800 infrequent Xable adjectives. After filtering for orthographical errors, brand names, and incorrect forms, a dataset of novel lexical items was isolated using the OED for evidence of attestation. Morphological patterns and usage contexts were analysed, highlighting factors such as (1) frequency distribution over 2010–2025, (2) geographical coverage, (3) contextual anchoring in the collocate environment, and 4) base form and derivative co-occurrence. Findings contribute insights into lexical innovation, morphological productivity, and the stages of lexical integration from a dynamic usage-based perspective (Schmid 2020).

Chris A. Smith · 0 citations
Open access Aug 2026

An improved metric for estimating morphological information in corpora

Abstract The emergence of large, consistently annotated corpora in many languages opens new avenues for linguistic typology by enabling the incorporation of usage-based evidence, including frequency information, into the study of language complexity. Morphological information is one of several facets of language complexity. In this paper, we use morphological feature annotations from the Universal Dependencies (UD) corpora to quantify information carried by morphology across 154 different datasets spanning 72 language varieties. We propose an information-theoretic approach that measures how surprising morphological feature values are given a token’s part of speech or lemma in the corpus. These token-level quantities are aggregated to dataset-level averages, yielding a usage-weighted estimate of morphological information load. We find substantial cross-linguistic variation in morphological information, and observe moderate to strong correlations with a related information-theoretic metric proposed by Çöltekin and Rama. By contrast, we find little correspondence with questionnaire-based typological metrics derived from Grambank, which represents an alternative approach to cross-linguistic comparison based on grammatical inventories rather than usage. This illustrates the difference between studying the possible extent of the grammatical system versus language use. We also discuss various drawbacks with corpus-based typology, such as comparability of datasets and uneven coverage across the globe.

Hedvig Skirgård, S. Mann · 0 citations
Jul 2026

Manual vs automated identification of English L2 suffixes and complex words across levels of proficiency

This report presents the preliminary evaluation of Morph , a web-based tool for the automatic counting of 51 English noun derivational suffixes in a controlled corpus of Mexican learners of English with proficiency levels from A2 to C1 on the Common European Framework of Reference for Languages (CEFR) scale. The evaluation consisted of a quantitative analysis of agreement between human annotators and Morph , a qualitative analysis to obtain the sources of disagreement, and finally, the evaluation of the tool’s efficiency using precision, recall, and F1 score metrics. Results displayed a high level of agreement between the human annotators and Morph . The main sources of disagreement were found to be human errors, learner misspellings, and automatic tagger issues. Finally, the efficiency metrics showed the tool to be effective in accurately identifying the target suffixes across proficiency levels, with a significant reduction in time and effort.

Ana Abigahil Flores-Hernández, Roilhi Frajo Ibarra-Hernández, P. Moore et al. · 0 citations
Open access Jul 2026

Low resource word sense disambiguation in Oromo with fine tuned small transformers.

A key task in natural language processing is word sense disambiguation (WSD), which attempts to determine the accurate meaning of ambiguous words based on their context. While transformer-based designs have achieved significant results in high-resource languages, WSD for low-resource languages such as Oromo remains hard due to inadequate annotated corpora and lexical resources. Contextual representation learning has been greatly enhanced by recent advancements in transformer-based language models, allowing for more reliable disambiguation in situations with limited input. This study uses a manually created dataset from the Oromo-English Dictionary to examine the efficacy of transformer-based models for lexical-sample WSD in Oromo. The dataset contains sentences annotated by two native speakers, attaining an inter-annotator agreement of 0.82, indicating good annotation reliability. The dataset was filtered for experimental usage following preprocessing, normalization, and elimination of noisy cases. 472 training sentences, 71 validation sentences (15%), and 140 test sentences made up the final dataset. The dataset has a highly unbalanced long-tail distribution and encompasses 43 sense classes. BERT-base-cased gets the best performance with an accuracy of 0.862 and a macro-F1 score of 0.2897, according to an experimental evaluation of transformer-based models, including BERT, RoBERTa, DistilBERT, Davlan/afro-xlmr-base, and multilingual variations. Significant differences between models with χ2 = 34.03 and p = 4.0 × 10-1 are confirmed by statistical analysis using the Friedman test. BERT-base-cased performs much better than most transformer variations and classical baselines, according to post-hoc Wilcoxon signed-rank tests. These results show that contextual transformer representations are quite successful for low-resource WSD, although there is still a significant class imbalance that limits performance.

Liyachew Edeti, Million Meshesha, Feda Negesse · 0 citations
Open access Jul 2026

Nominalization Patterns in Tombatu Language: A Generative Morphological Analysis of Affixal Noun Formation

Regional languages preserve complex grammatical systems that reveal how communities organize experience, identity, and cultural knowledge. However, nominalization in Tombatu, an underdescribed Austronesian language of Minahasa, has not been systematically examined through an explicit rule-based morphological framework. This study addresses this gap by investigating the structural patterns, semantic functions, and generative mechanisms of affixal noun formation in Tombatu. Employing a descriptive-taxonomic design, data were collected over six months in Tombatu Village from proficient speakers through naturally occurring speech, elicitation, recording, note-taking, and informant validation. The final dataset found unique nominalized forms. These forms were classified into three major processes comprising 11 noun-forming prefixes, one noun-forming suffix, and 14 affix combinations, yielding 26 identified affixal configurations. The analysis applied Word Combination Rules, Derivation Rules, the Boundary Insertion Convention, and relevant morphological constraints. The findings show that prefixation exhibits the greatest structural and semantic variation in the dataset, producing agentive, person-denoting, instrumental, object-related, and result-related nouns. For example, /tapə-/ combines with /lukuʔ/ ‘to drink’ to form /tapəlukuʔ/ [tapəlukuʔ] ‘drinker’. Suffixation is more restricted, with /-an/ forming locative nouns, as in /tawoi/ ‘to work’ plus /-an/, which produces /tawoian/ [tawoian] ‘workplace’. Complex affix combinations form nouns denoting places, processes, objects, states, results, and entities. This study provides the first systematic generative account of Tombatu nominalization, expands empirical evidence for Austronesian word formation, strengthens regional-language documentation, and offers a linguistic basis for language maintenance and potential contrastive morphological activities in multilingual education.

Verra E. Manangkot, J. Liuw, L. Amalputra et al. · 0 citations