Skip to content
Open access

CLASSIFICATION OF RUSSIAN-LANGUAGE SHORT TEXTS: COMPARISON OF TF-IDF+ML AND SIMPLE NEURAL NETWORK MODELS BASED ON UNIFIED ASSESSMENT PROTOCOLS

J. J. Raxmani
Aug 2026 · ИНФОРМАЦИОННЫЕ СИСТЕМЫ И ТЕХНОЛОГИИ · Vol 25, pp. 33-41 · 0 citations

TL;DR

The results show that with a limited amount of training data, TF-IDF-based models provide quality comparable to simple neural network architectures, while neural network approaches demonstrate an advantage when increasing the size of the corpus.

Abstract

The paper considers the problem of automatic classification of Russian-language short texts using traditional statistical and neural network approaches. The aim of the study is to compare the effectiveness of TF-IDF models in combination with classical machine learning algorithms (SVM, logistic regression) and a simple neural network architecture when solving the same classification problem using unified assessment protocols. During the experiments, the influence of various methods of text preprocessing (cleaning, stemming, lemmatization) on the classification quality and the stability of models to noisy data is analyzed. The novelty of the work lies in a replicated comparison of traditional and neural network methods on Russian-language corpora, as well as in assessing the impact of morphological normalization on quality indicators. The results show that with a limited amount of training data, TF-IDF-based models provide quality comparable to simple neural network architectures, while neural network approaches demonstrate an advantage when increasing the size of the corpus. The results of the research can be used in the development of text analysis systems, content filtering and intelligent dialog interfaces for the Russian language.

Read PDF

Similar papers

Review Open access Aug 2026

EXPERIMENTAL COMPARISON OF TEXT VECTORIZATION METHODS FOR SENTIMENT ANALYSIS TASK

The paper presents a comparative analysis of the effectiveness of various text vectorization methods for the task of Sentiment Analysis of Russian-language reviews. The study covers classical frequency-based approaches (TF IDF, n-grams), statistical models (Word2Vec, FastText), and a contextual method based on the pre-...

O. I. Zakharova, S. Bednyak, Yaroslav Dmitrievich Kanunnikov · 0 citations
Open access Jul 2026

Implementation of a Bi-LSTM Model for Automatic Text Classification of Mathematics, Science, and Indonesian Language Questions

It can be concluded that the Bi-LSTM model is effective for automatic text classification of educational questions and has strong potential for further development in technology-based question grouping systems.

Mochamad soffan Muslim, Aviv Yuniar Rahman, R. Pahlevi · 0 citations
Open access 2026

Comparative Analysis of Language Models for Sentiment Classification

Comparing and analysing the performance of several machine learning algorithms on fine-grained sentiment classification problems to examine their suitability and shortcomings for use as models in sentiment analysis suggests large language models perform significantly worse on the 28-class classification task in zero-sh...

Shangjiafeng Guo · 0 citations
#small language model Open access Aug 2026

Incorporating Linguistic Normalization in Croatian NLP: Evaluating the Impact of Lemmatization on Disinformation Detection Performance

It is demonstrated that lemmatization does not produce uniform gains across architectures: while linear models and croBERT display small but measurable improvements from morphological normalization, non-linear models such as RBF SVM and neural networks experience substantial declines in performance.

I. Ljubi, M. Horvat, G. Gledec et al. · 0 citations
Review Open access Aug 2026

OPTIMIZING MACHINE LEARNING CLASSIFIERS FOR HIGH-ACCURACY SENTIMENT DETECTION IN THE TURKISH LANGUAGE

This study investigates the optimization of classical machine learning classifiers and ensemble learning strategies for binary Turkish sentiment analysis under a unified experimental framework and demonstrates that optimized classical models remain highly effective in Turkish SA, and their accuracy can be further impro...

Ahmad Bwidani, Ali A. H. Karah Bash · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.