Skip to content
Open access

Characterization and Mechanisms of Lexical Complexity in AI-Generated Texts: A Comparative Corpus-Based Study

2026 · Lecture Notes on Language and Literature · 0 citations · 15 references

TL;DR

Analysis of AIGC texts points out that the complexity of AI text primarily stems from its mechanism of selecting vocabulary based on probability distributions, which favors longer words, abstract nouns, and words with high semantic content, thereby forming a highly compact linguistic surface.

Abstract

: Based on a corpus-based methodology, this study analyzes the intrinsic reasons for the high level of lexical complexity observed in Artificial Intelligence Generated Content (AIGC). The research compares 24 English argumentative essays written by AI with 24 second-language (L2) learner essays reaching the IELTS Writing Task 2 Band 7 level. Under controlled conditions of identical genre and topic, the study performs quantitative statistics across three dimensions: lexical sophistication, semantic abstraction, and information density. Statistical results indicate that the frequency of advanced vocabulary in AI texts is significantly higher, approximately 2.3 times that of human texts. The proportion of abstract nouns reached 9.14%, far exceeding the 3.32% found in human texts, suggesting that AI expressions tend toward nominalization and conceptualization. Regarding overall information organization, the lexical density of AI texts was 69.9%, also surpassing the 60.3% of human texts, reflecting a stronger tendency for information condensation and phrasal structures. The analysis points out that the complexity of AI text primarily stems from its mechanism of selecting vocabulary based on probability distributions. This mechanism favors longer words, abstract nouns, and words with high semantic content, thereby forming a highly compact linguistic surface. Such complexity is essentially a formal feature at the statistical level and is not entirely equivalent to the proficiency levels corresponding to human L2 acquisition. These findings provide empirical references for AI text identification, the refinement of writing evaluation standards, and L2 writing pedagogy.

Read PDF

Similar papers

Open access Jul 2026

Syntactic Complexity in AI-Generated vs. Human-Authored Linguistic and Literary Texts

The results indicate that the complexity of syntax is genre-based and not source-based and in general, the discipline genre had a more significant effect on syntax variation than the authorship source.

Asia A. Alheety, Meethaq Khamees Khalaf, H. Mohammed · 0 citations
Open access Jul 2026

Texts Generated by Artificial Intelligence: Structure and Semantics

It was concluded that texts generated by artificial intelligence constitute a separate linguistic phenomenon with its own set of characteristics, which requires a special typology and a flexible, updatable analysis methodology.

L. Kravets, Viktória Stefuca, N. Libak et al. · 0 citations
Open access Jul 2026

Human vs AI-Generated Texts in Language Learning: A Linguistic Comparison

The findings show that AI-generated texts exhibit greater lexical diversity and syntactic complexity; however, they often exhibit structural uniformity, overuse of cohesive devices, and limited pragmatic depth, and should not replace professionally designed educational materials.

V. Smaglii, T. Korolova, Svitlana Yukhymets et al. · 0 citations
Open access Jul 2026

FROM CLAUSES TO COGNITION: A COMPUTATIONAL ANALYSIS OF LINGUISTIC COMPLEXITY IN INTERMEDIATE BOOK 2

This study investigates the syntactic complexity of Intermediate English Book 2 within the Pakistani curriculum through a corpus-based and computational approach. Drawing on methods from Natural Language Processing, the analysis employs established syntactic indices—mean length of sentence (MLS), mean length of clause (MLC), clauses per sentence (C/S), and dependent clauses per T-unit (DC/TU)—to examine structural patterns across lessons and text types. The findings reveal that the textbook exhibits moderate to high linguistic complexity, with MLS values ranging approximately between 17 and 20 and consistent subordination levels (DC/TU ≈ 0.4), aligning with B2–C1 proficiency benchmarks. The results further demonstrate that complexity is not uniform but varies across genres: scientific and historical texts show higher syntactic density and subordination, while narrative and humorous texts rely relatively more on coordination and linear structures. Clause-level analysis indicates a strong presence of dependent clauses, particularly adverbial and relative clauses, which function to encode causal relationships, temporal sequencing, and descriptive detail. From a cognitive perspective, these features contribute to increased processing demands, requiring learners to engage with hierarchically structured information. The study argues that syntactic complexity in Book 2 functions as a cognitively demanding yet pedagogically purposeful feature, supporting the transition from intermediate to advanced proficiency. By integrating NLP-based analysis with SLA theory, the research provides an empirical framework for evaluating textbook difficulty and highlights the need for scaffolded instruction to manage cognitive load. The findings have implications for learners, educators, and curriculum planners, emphasizing the importance of balancing linguistic richness with accessibility in instructional materials.

Azhar Munir Bhatti, Prof. Dr. Ahsan Bashir · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.