Skip to content
Open access

PERPLEXITY-BASED AI-GENERATED TEXT CLASSIFICATION IN UKRAINIAN USING SMALL LANGUAGE MODELS

Jul 2026 · Advanced Information Systems · Vol 10, pp. 5-12 · 0 citations

TL;DR

The comparative analysis with a traditional NLP-based discriminative neural network model revealed that direct text piece classification outperforms perplexity-based methods, although the latter still demonstrate practical utility.

Abstract

The aim of the research. The rapid advancement of generative artificial intelligence language models has introduced new complexities in discerning the authorship and quality of textual content. In this paper, we explored the feasibility of using perplexity – a measure of token predictability – as the only discriminative feature for classifying AI-generated versus human-written texts in Ukrainian within the IT domain. Our approach employed small language models to calculate perplexity and detect content generated by state-of-the-art models, evaluating the potential for lightweight solutions. Research results. Initial experiments using a single perplexity threshold across Gemma 3 / Llama 3.2 1B models yielded classification accuracies around 0.70. The full token-level probability sequences were proposed as feature vectors, enabling us to achieve an accuracy of 0.68 via simple KNN classification. Finally, the convolutional neural network architectures trained on these features allowed us to obtain 0.82–0.87 accuracy. Conclusions. The comparative analysis with a traditional NLP-based discriminative neural network model revealed that direct text piece classification outperforms perplexity-based methods, although the latter still demonstrate practical utility.

Read PDF

Similar papers

Open access Aug 2026

Text Classification Using Large Language Models to Detect AI and Human-Generated Sentences

With the rapid development of the technology of Artificial Intelligence (AI), especially Large Language Models (LLMs) like ChatGPT, GPT-4 and Gemini, systems have become able to generate texts very similar to human writing. This similarity has been a boon to various sectors, but also poses fresh challenges of content a...

Arfan Rahman Yudiantoro, Prati Hutari Gani, Donni Richasdy · 0 citations
Open access Aug 2026

Arabic Plagiarism Detection Using Word2Vec-Based Semantic Features and Random Forest Classification on the ExAraPlagDet Dataset

The findings underscore the potential of advanced NLP techniques to overcome language-specific challenges, providing a foundation for future research in multilingual plagiarism detection and enhancing the development of tools for other languages facing similar challenges.

Hanan Fawzy, Ahmad Salah, Heba El-Fiqi et al. · 0 citations
Open access Aug 2026

Detecting AI-Generated Bulgarian Text: A Two-Step Multi-Class Classification Approach

This paper introduces a novel two-step multi-class classification system to identify varying degrees of machine involvement in Bulgarian text. As Large Language Models (LLMs) proliferate, distinguishing original human writing from machine-assisted or machine-generated content is crucial to prevent misinformation and pr...

Boyan Bogdanov, D. Georgiev, D. Dimitrov et al. · 0 citations
Open access Jul 2026

LSTM-Based Classification of Indonesian Regional Song Lyrics by Language

This study successfully proposes a Long Short-Term Memory (LSTM)-based model for automatic classification of Indonesian regional song lyrics by language, demonstrating that LSTM effectively captures sequential linguistic patterns and contextual relationships within regional languages.

Muhammad Rizky, Anandita Priatama, Aviv Yuniar Rahman et al. · 0 citations
Review Open access Aug 2026

Mapping the Landscape of AI-driven Quranic Recitation Recognition: A Bibliometric and Thematic Analysis (2016–2026)

This study aims to explore the development and application of artificial intelligence techniques in Arabic speech recognition, with a specific focus on recitation accuracy, pronunciation analysis, and language learning support. It seeks to identify trends, methods, and challenges in AI-based Arabic recitation recogniti...

Khirulnizam Abd Rahman, Che Wan Shamsul Bahri Che Wan Ahmad, M. Lubis et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.