Hybrid deep learning for three-way classification of human-written, AI-generated, and AI-rephrased text
Abstract
Abstract The rapid proliferation of large language models (LLMs) has made it increasingly difficult to distinguish AI-generated and AI-rephrased text from authentic human writing, posing serious risks to academic integrity, journalism, and e-commerce credibility. This paper presents a rigorous benchmark evaluation for three-way AI text origin classification designed to detect human-written, AI-generated, and AI-rephrased text by integrating handcrafted linguistic features, transformer-based contextual embeddings, and ensemble learning strategies. A corpus of 161,788 samples is curated from Amazon product reviews and Twitter posts, with AI-generated text produced via zero-shot prompting and AI-rephrased text produced via few-shot prompting using GPT, DeepSeek, and Kimi. The corpus comprises 46,844 human-written, 57,247 AI-generated, and 57,697 AI-rephrased samples distributed across training (113,249), validation (16,180), and test (32,359) splits. A hybrid feature vector of 168 dimensions is constructed, combining 68 handcrafted stylistic, lexical, linguistic, syntactic, and stylometric features with 100-dimensional embeddings derived from RoBERTa and the all-MiniLM-L6-v2 Sentence Transformer. Classical machine learning models, Random Forest, SVM, and Logistic Regression, are trained on this hybrid representation, while RoBERTa-base and DeBERTa-v3-base are fine-tuned end-to-end on raw tokenized text. Experimental evaluation demonstrates that transformer models substantially outperform classical methods, with RoBERTa achieving 95.49% accuracy and DeBERTa-v3 achieving 94.88%. A probability-averaging ensemble further improves accuracy to 95.66%, and a stacking ensemble with a Random Forest meta-learner reaches 96.46%. Ablation studies confirm the complementary value of handcrafted features alongside contextual embeddings. Across all models, AI-rephrased content is the most challenging category due to its semantic proximity to human writing, highlighting a persistent frontier for future research. The dataset and evaluation framework provide a reproducible foundation for future AI content detection research, with generalization to unseen generators and domains identified as primary directions for future work.