Skip to content

Author

Y. Karabaliyev

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Morpheme-Aware Interpolated N-gram Language Modeling for Low-Resource Speech Recognition

This paper presents KazMorphLM, a morpheme-aware language model for Kazakh automatic speech recognition (ASR). Kazakh, a highly agglutinative Turkic language, poses a fundamental challenge for conventional word-level language models, since a single root can generate hundreds of inflected forms through productive suffixation, causing extreme data sparsity. Our objective is to overcome this sparsity by modelling language at the morpheme level. The method combines three components: (1) a rule-based morpheme segmenter built on a fully categorized inventory of 118 suffix entries (175 surface forms) across 12 morphological categories, with vowel-harmony validation and consonant-assimilation rules; (2) a two-level interpolated n-gram architecture coupling a 7-gram morpheme model with a 5-gram word model under Witten-Bell smoothing; and (3) a four-channel rescoring mechanism integrating acoustic, word-level, morpheme-level and vowel-harmony scores. Integrated into a hybrid FastConformer–MMS-1B pipeline, KazMorphLM attains 6.86% word error rate (WER) on the FLEURS test set under N-best rescoring, a 15.2% relative reduction over word-level KenLM rescoring (8.09%). On a live evaluation with 44 native speakers (421 recordings), KazMorphLM significantly outperforms word-level KenLM rescoring (6.01% vs. 7.45% WER; Wilcoxon p<0.001). To our knowledge this is the first morpheme-aware, vowel-harmony-informed language model for Kazakh ASR rescoring, with a methodology transferable to other Turkic languages.

Y. Karabaliyev, Kateryna Kolesnikova, Khlevnaya Yulia · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.