Jul 2026· Advanced Information Systems· Vol 10, pp. 5-12· 0 citations
TL;DR
The comparative analysis with a traditional NLP-based discriminative neural network model revealed that direct text piece classification outperforms perplexity-based methods, although the latter still demonstrate practical utility.
Abstract
The aim of the research. The rapid advancement of generative artificial intelligence language models has introduced new complexities in discerning the authorship and quality of textual content. In this paper, we explored the feasibility of using perplexity – a measure of token predictability – as the only discriminative feature for classifying AI-generated versus human-written texts in Ukrainian within the IT domain. Our approach employed small language models to calculate perplexity and detect content generated by state-of-the-art models, evaluating the potential for lightweight solutions. Research results. Initial experiments using a single perplexity threshold across Gemma 3 / Llama 3.2 1B models yielded classification accuracies around 0.70. The full token-level probability sequences were proposed as feature vectors, enabling us to achieve an accuracy of 0.68 via simple KNN classification. Finally, the convolutional neural network architectures trained on these features allowed us to obtain 0.82–0.87 accuracy. Conclusions. The comparative analysis with a traditional NLP-based discriminative neural network model revealed that direct text piece classification outperforms perplexity-based methods, although the latter still demonstrate practical utility.
With the rapid development of the technology of Artificial Intelligence (AI), especially Large Language Models (LLMs) like ChatGPT, GPT-4 and Gemini, systems have become able to generate texts very similar to human writing. This similarity has been a boon to various sectors, but also poses fresh challenges of content a...
Arfan Rahman Yudiantoro, Prati Hutari Gani, Donni Richasdy· International Journal on Inf...· 0 citations
The findings underscore the potential of advanced NLP techniques to overcome language-specific challenges, providing a foundation for future research in multilingual plagiarism detection and enhancing the development of tools for other languages facing similar challenges.
Hanan Fawzy, Ahmad Salah, Heba El-Fiqi et al.· Informatica· 0 citations
This paper introduces a novel two-step multi-class classification system to identify varying degrees of machine involvement in Bulgarian text. As Large Language Models (LLMs) proliferate, distinguishing original human writing from machine-assisted or machine-generated content is crucial to prevent misinformation and pr...
Boyan Bogdanov, D. Georgiev, D. Dimitrov et al.· Computational Linguistics in...· 0 citations
This study successfully proposes a Long Short-Term Memory (LSTM)-based model for automatic classification of Indonesian regional song lyrics by language, demonstrating that LSTM effectively captures sequential linguistic patterns and contextual relationships within regional languages.
Muhammad Rizky, Anandita Priatama, Aviv Yuniar Rahman et al.· Buana Information Technology...· 0 citations
This study aims to explore the development and application of artificial intelligence techniques in Arabic speech recognition, with a specific focus on recitation accuracy, pronunciation analysis, and language learning support. It seeks to identify trends, methods, and challenges in AI-based Arabic recitation recogniti...
Khirulnizam Abd Rahman, Che Wan Shamsul Bahri Che Wan Ahmad, M. Lubis et al.· e-Jurnal Penyelidikan dan In...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.