MORPHOLOGICAL DISAMBIGUATION FOR THE KAZAKH LANGUAGE USING TRANSFORMER-BASED MODELS
Abstract
Morphological ambiguity constitutes a significant challenge for natural language processing in agglutinative languages, as a single word form might include many grammatical categories. The Kazakh language features productive suffixation, vowel harmony, and intricate morphophonological patterns, which considerably hinder automatic morphological analysis. This work presents a transformer-based methodology for morphological disambiguation in Kazakh texts, with the objective of identifying the appropriate morphological interpretation of word forms within context. A contextual language model tailored for Kazakh is refined for token-level morphological tagging utilizing a manually validated annotated corpus of news articles. The suggested method utilizes self-attention mechanisms to capture long-range contextual dependencies that are challenging to represent with conventional rulebased or recurrent neural techniques. The experimental assessment reveals that the transformer-based model attains superior accuracy and F1-score relative to rule-based morphological analyzers and BiLSTM-based benchmarks. The findings demonstrate that contextualized embeddings significantly enhance the resolution of morphological ambiguity, especially with homonymous suffixes and infrequent grammatical structures. The results validate the efficacy of transformer topologies for low-resource agglutinative languages and establish a feasible basis for incorporating morphology-aware models into comprehensive Kazakh natural language processing frameworks.