Problems of Artificial Intelligence for the Azerbaijani Language: The Impact of Limited Language Corpus
Artificial intelligence (AI) systems, such as large-scale speech recognition systems, are vital for acquiring linguistic competence and achieving this is essential for their use. Thus, for large-scale speech recognition systems, acquisition of linguistic competence is essential for usage of the system. Even for the widely spoken language like English, Mandarin and Spanish, these corpora are vast and have hundreds of billions of tokens, which can be used to power up virtually all natural-language processing (NLP) tasks. In the online world, however, the Azerbaijani language is far less well resourced: publicly available language corpora are limited to a small number of high-resource languages. In the article, the main problems related to corpus scarcity, such as poor quality of machine translation, incorrect morphological processing, recognition dialects of the Azerbaijan language and failure of chatbots to understand the Azerbaijan language are discussed.