Detecting AI-Generated Essays with Lexical Stylometric Features using Machine Learning
The increasing use of generative artificial intelligence (AI) in academic writing has intensified concerns about authorship authenticity. This study investigates whether lexical stylometric features alone can distinguish AI-generated essays from human-written essays. A balanced dataset of 200 undergraduate essays (100 AI-generated and 100 human-written) was analyzed using vocabulary richness measures, lexical density, word usage distributions, and word-level n-grams. Logistic Regression, Decision Tree, and Support Vector Machine classifiers were applied to evaluate performance. Results indicate that lexical features provide strong discriminatory capability across the three models with accuracy ranging between 90 – 97.5%, suggesting that effective AI authorship detection can be achieved without complex feature engineering or ensemble methods. The findings support the feasibility of a lightweight and interpretable detection framework for educational contexts.