1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

An improved metric for estimating morphological information in corpora

Abstract The emergence of large, consistently annotated corpora in many languages opens new avenues for linguistic typology by enabling the incorporation of usage-based evidence, including frequency information, into the study of language complexity. Morphological information is one of several facets of language complexity. In this paper, we use morphological feature annotations from the Universal Dependencies (UD) corpora to quantify information carried by morphology across 154 different datasets spanning 72 language varieties. We propose an information-theoretic approach that measures how surprising morphological feature values are given a token’s part of speech or lemma in the corpus. These token-level quantities are aggregated to dataset-level averages, yielding a usage-weighted estimate of morphological information load. We find substantial cross-linguistic variation in morphological information, and observe moderate to strong correlations with a related information-theoretic metric proposed by Çöltekin and Rama. By contrast, we find little correspondence with questionnaire-based typological metrics derived from Grambank, which represents an alternative approach to cross-linguistic comparison based on grammatical inventories rather than usage. This illustrates the difference between studying the possible extent of the grammatical system versus language use. We also discuss various drawbacks with corpus-based typology, such as comparability of datasets and uneven coverage across the globe.

Hedvig Skirgård, S. Mann · 0 citations