CLASSIFICATION OF RUSSIAN-LANGUAGE SHORT TEXTS: COMPARISON OF TF-IDF+ML AND SIMPLE NEURAL NETWORK MODELS BASED ON UNIFIED ASSESSMENT PROTOCOLS
J. J. Raxmani
Aug 2026· ИНФОРМАЦИОННЫЕ СИСТЕМЫ И ТЕХНОЛОГИИ· Vol 25, pp. 33-41· 0 citations
TL;DR
The results show that with a limited amount of training data, TF-IDF-based models provide quality comparable to simple neural network architectures, while neural network approaches demonstrate an advantage when increasing the size of the corpus.
Abstract
The paper considers the problem of automatic classification of Russian-language short texts using traditional statistical and neural network approaches. The aim of the study is to compare the effectiveness of TF-IDF models in combination with classical machine learning algorithms (SVM, logistic regression) and a simple neural network architecture when solving the same classification problem using unified assessment protocols. During the experiments, the influence of various methods of text preprocessing (cleaning, stemming, lemmatization) on the classification quality and the stability of models to noisy data is analyzed.
The novelty of the work lies in a replicated comparison of traditional and neural network methods on Russian-language corpora, as well as in assessing the impact of morphological normalization on quality indicators. The results show that with a limited amount of training data, TF-IDF-based models provide quality comparable to simple neural network architectures, while neural network approaches demonstrate an advantage when increasing the size of the corpus.
The results of the research can be used in the development of text analysis systems, content filtering and intelligent dialog interfaces for the Russian language.
The paper presents a comparative analysis of the effectiveness of various text vectorization methods for the task of Sentiment Analysis of Russian-language reviews. The study covers classical frequency-based approaches (TF IDF, n-grams), statistical models (Word2Vec, FastText), and a contextual method based on the pre-...
O. I. Zakharova, S. Bednyak, Yaroslav Dmitrievich Kanunnikov· Infokommunikacionnye tehnolo...· 0 citations
It can be concluded that the Bi-LSTM model is effective for automatic text classification of educational questions and has strong potential for further development in technology-based question grouping systems.
Mochamad soffan Muslim, Aviv Yuniar Rahman, R. Pahlevi· Buana Information Technology...· 0 citations
Comparing and analysing the performance of several machine learning algorithms on fine-grained sentiment classification problems to examine their suitability and shortcomings for use as models in sentiment analysis suggests large language models perform significantly worse on the 28-class classification task in zero-sh...
Shangjiafeng Guo· International journal of eng...· 0 citations
It is demonstrated that lemmatization does not produce uniform gains across architectures: while linear models and croBERT display small but measurable improvements from morphological normalization, non-linear models such as RBF SVM and neural networks experience substantial declines in performance.
I. Ljubi, M. Horvat, G. Gledec et al.· Electronics· 0 citations
This study investigates the optimization of classical machine learning classifiers and ensemble learning strategies for binary Turkish sentiment analysis under a unified experimental framework and demonstrates that optimized classical models remain highly effective in Turkish SA, and their accuracy can be further impro...
Ahmad Bwidani, Ali A. H. Karah Bash· Uludağ University Journal of...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.