Skip to content
Open access

Multilingual AI-Generated Text Detection in Arabic, English, and Turkish Using a Hybrid Transformer–Graph Convolutional Network

Jul 2026 · Applied Sciences · 0 citations · 32 references

TL;DR

A hybrid architecture that combines a Transformer-based DistilBERT model with a Graph Convolutional Network (GCN) that enhances detection by modeling structural relationships within text data is proposed.

Abstract

Detecting AI-generated text has become a critical task as artificial intelligence systems are increasingly used in content creation. Current detection methods often suffer from limited accuracy and weak multilingual performance. This problem is especially challenging in Turkish, Arabic, and English due to their distinct linguistic structures, including agglutinative morphology in Turkish, root-based morphology in Arabic, and semantic ambiguity in English. To address these challenges, this study proposes a hybrid architecture that combines a Transformer-based DistilBERT model with a Graph Convolutional Network (GCN). While DistilBERT captures rich contextual and semantic information, GCN enhances detection by modeling structural relationships within text data. The proposed model is evaluated against other well-known approaches. Experimental results show that the hybrid DistilBERTGCN framework achieves high detection accuracy, reaching 99% for English and 98% for Turkish and Arabic. In addition, this study introduces new multilingual datasets, contributing to the advancement of the literature research.

Read PDF

Similar papers

Open access Aug 2026

An intelligent deep learning for Arabic stemming and morphological classification

Arabic language, due to its complex morphology and richness of grammar features poses significant challenges in natural language processing (NLP). In this paper, we propose a two-stage deep learning pipeline that combines Arabic text stemming and morphological classification within a single deep learning architecture. The relationship between morphological reduction and grammatical categorization is exploited by combining character-level sequence processing with transformer-based classification. A bidirectional long short-term memory (Bi-LSTM) model is employed for Arabic stem extraction to build a sequence-to-sequence (seq2seq) stemming model named Char Stemmer. To evaluate the proposed model, a gold standard dataset consisting of 260,000 traditional Arabic words extracted from Quranic words and classical Arabic books is utilized. This dataset contains a wide range of challenging word structures suitable for robust evaluation. The Char Stemmer achieved an accuracy of 93.88% on the stemming task. The proposed model obtained 93.88% accuracy, demonstrating a 38% improvement over the best traditional stemmer, P-Stemmer. Beyond stemming, the impact of stemmers on subsequent tasks is evaluated, particularly Arabic word classification. Words are categorized into three morphological classes: noun, verb, and particle. Experimental results show that the proposed system achieved macro average precision, recall, and F1-score of 0.91, 0.89, and 0.90, respectively, with an overall classification accuracy of approximately 99%.

Azal Alaswaad, B. Minaei-Bidgoli · 0 citations
Open access 2026

Hybrid Lexical–Contextual Learning for Arabic News Classification: Integrating TF–IDF with AraBERT Embeddings

Arabic text classification is still a daunting undertaking because of the rich morphology, derivational complexity as well as lexical variability that is inherent in the language. Though transformer-based pretrained models have been a major breakthrough in Arabic Natural Language Processing (NLP), recent data indicates that conventional lexical representations continue to give a good discriminating ability in structured tasks like news classification. The given study will be a comparative analysis of classical machine learning models, transformer-based fine-tuning, and lexical contextual fusion framework as a hybrid method in Arabic multi-class news classification. A subset of the MAAD dataset (13,866 Arabic news articles) in six categories was carefully selected and cleaned and sampled to conduct experiments on a balanced and well-cleaned subset of the entire data set. To guarantee the data quality, the Arabic-character ratio filtering and length constraints were used to eliminate the corrupted and non-Arabic samples. TF-IDF features with Logistic regression and Linear Support Vector machine (SVM) were used to implement baseline models and contextual modeling was done using fine-tuning AraBERTv2. Moreover, the hybrid feature-level fusion model was suggested by adding TFIDF vectors with contextual embeddings were obtained using AraBERT. Experimental findings indicate that the traditional linear models are still very competitive whereby Linear SVM has an accuracy of 95.57 percent. The fine-tuned AraBERT obtained an accuracy of 91.6%, which puts lexical features in structured news datasets in the spotlight as still significant. The hybrid model with the highest performance in the proposed structure was 96.03% accurate and 96.02% Macro-F1 score, which shows that combining lexical statistical cues and contextual embeddings is complementary. These results show that hybrid lexical-contextual representations offer a powerful, computationally effective solution to Arabic news classification. The analysis offers reproducible experimental environments and intricate statistical investigation, to add the empirical data regarding the interaction between classical and deep methods of learning Arabic NLP.

Y. Farhan, Mustafa Tareq, Boumedyen Shannaq et al. · 0 citations
Open access 2026

Arabic News Text Classification Using Deep Learning Models with Dynamic N-grams

The complexity and morphological richness of the Arabic language pose significant challenges in natural language processing (NLP), including issues with contextual understanding and feature extraction. Traditional deep learning architectures such as CNNs, LSTMs, and GRUs often struggle to effectively model these linguistic intricacies, limiting their performance on Arabic text analysis tasks. To address these limitations, the study integrates parallel multi-kernel word-level convolutional features into conventional and hybrid deep learning models. The convolutional windows capture short contextual relationships among neighboring word tokens, while recurrent components model longer sequential evidence; the framework does not directly analyze roots, affixes, or other within-word morphological structures. These enhancements are integrated into CNN, LSTM, GRU, and hybrid architectures such as LSTM-CNN and GRU-CNN. A comprehensive evaluation was conducted across varying learning rates to assess the impact of the enhanced configurations on model performance. The results indicate competitive performance within the evaluated architectures and dataset variants, although the magnitude of improvement depends on the model and learning rate. Under their best settings, the Dynamic N-gram LSTM-CNN achieved an accuracy of 93.32%, while the Dynamic N-gram LSTM achieved 93.57%. Because previously published studies use different corpora, class configurations, preprocessing pipelines, and evaluation protocols, these results are not presented as evidence of state-of-the-art superiority. Instead, the study provides a systematic within-study assessment of model sensitivity to architecture, preprocessing, and learning-rate selection. Future directions include character- and subword-level modeling, transformer-based architectures, and domain-specific tasks such as sentiment analysis and information retrieval.

Ahmed I.Taloba, George Samy Rady, Khaled F. Hussain · 0 citations
Open access Jul 2026

Implementation of the BiLSTM Model for Detecting AI-Generated Indonesian Text

The rapid advancement of generative Artificial Intelligence (AI) presents challenges to academic integrity due to potential misuse like plagiarism. This study develops a text detection system specifically for the Indonesian language using a Deep Learning approach with a Bidirectional Long Short-Term Memory (Bi-LSTM) architecture. The research methodology follows the Cross-Industry Standard Process for Data Mining (CRISP-DM) framework. A dataset comprising 5,008 text rows was compiled via web scraping from journalism platforms and academic journals indexed in SINTA 4 for human-written texts, while AI-generated counterparts were engineered using ChatGPT and Google Gemini paraphrases. Text features were extracted using a Keras Tokenizer and Embedding Layer with 64 dimensions. Evaluation of the trained Bi-LSTM model on a 30% validation split demonstrated an overall accuracy of 78.24% and a Mean Absolute Error (MAE) of 0.3295. Specifically, the model achieved a 93.77% success rate in identifying human-written texts, though it logged a lower detection rate of 62.62% for academic AI text structures. The final model was successfully deployed as a web application using Streamlit.

Rafil Moehamad Alif, Syariful Alam, Chandra Dewi Lestari · 0 citations
Open access Aug 2026

Arabic Plagiarism Detection Using Word2Vec-Based Semantic Features and Random Forest Classification on the ExAraPlagDet Dataset

The findings underscore the potential of advanced NLP techniques to overcome language-specific challenges, providing a foundation for future research in multilingual plagiarism detection and enhancing the development of tools for other languages facing similar challenges.

Hanan Mohammed Fawzy, Ahmad Salah, Heba El-Fiqi et al. · 0 citations
Conference Jul 2026

A Comparative Benchmark of Specialized Deep Learning Architectures and Fine-Tuned LLMs for Arabic Text Readability

Readability assessment for Arabic remains challenging due to the language's complex morphology. This paper presents a comparative study benchmarking traditional Machine Learning (ML), advanced Deep Learning (DL), and finetuned Large Language Models (LLMs). Utilizing a dataset of 4,519 Arabic sentences categorized into three proficiency levels, we evaluate models across accuracy and computational efficiency. Our results demonstrate that a hybrid CNN-BiLSTM architecture utilizing AraVec (Word2Vec) embeddings achieves a peak accuracy of 96.68%, outperforming fine-tuned LLMs like Llama3.2-1B (94.69%). We provide empirical evidence of the prohibitive resource demands in LLMs, which required significantly higher training times (14,697s) compared to specialized DL models (162.98). These findings suggest that for discrete Arabic text classification, tailored DL architectures provide a superior balance of precision and resource efficiency.

Mohamed-Amine Ouassil, Rabia Rachidi, Othmane Daanouni et al. · 0 citations