Skip to content
Open access

Syntactic Complexity in AI-Generated vs. Human-Authored Linguistic and Literary Texts

Jul 2026 · Arab World English Journal · 0 citations · 13 references

TL;DR

The results indicate that the complexity of syntax is genre-based and not source-based and in general, the discipline genre had a more significant effect on syntax variation than the authorship source.

Abstract

This paper examines how the syntactic complexity of academic writing is affected mainly by the source of authorship (AI-generated or human) or by the genre of disciplinary writing (linguistic or literary). The primary purpose is to test the syntactic-complexity differences between these variables and to establish the degree of influence of genre conventions on structural variation. The importance of the research is that it adds to the existing discussions about AI as a phenomenon in academic writing and, specifically, whether AI-generated texts are capable of syntactically reproducing the specific norms of a specific discipline of writing. To this end, the comparative corpus-based design was used. The sample consisted of 20 introduction sections: equal numbers of linguistic and literary texts and equal numbers of human-authored AI-generated texts. The Second Language Syntactic Complexity Analyzer (L2SCA) extracts fourteen syntactic complexity measures, which include length of production unit, subordination, coordination and phrasal sophistication. The results indicate that the complexity of syntax is genre-based and not source-based. Although no differences were found to be constant in both AI-generated and human academic introductions in linguistic data, there were much higher levels of subordination in literary texts of human origin. In general, the discipline genre had a more significant effect on syntax variation than the authorship source. The paper suggests the implementation of genre-sensitive models to assess AI-written academic texts and recommends additional studies that would use a bigger sample and discourse analysis.

Read PDF

Similar papers

Open access 2026

Characterization and Mechanisms of Lexical Complexity in AI-Generated Texts: A Comparative Corpus-Based Study

Analysis of AIGC texts points out that the complexity of AI text primarily stems from its mechanism of selecting vocabulary based on probability distributions, which favors longer words, abstract nouns, and words with high semantic content, thereby forming a highly compact linguistic surface.

Lulu Chen · 0 citations
Review Open access Jul 2026

Linguistic Features of AI-Generated Academic Texts and the Role of Human Editing

A statistically significant increase in lexical density and frequency of cohesive markers has been revealed, indicating an increase in information compression and explicit discursive organization of texts, consistent with characteristics described in AI-assisted writing studies.

T. Nedashkivska, I. Varvaruk, M. Podoliak et al. · 0 citations
Open access Jul 2026

Human vs AI-Generated Texts in Language Learning: A Linguistic Comparison

The findings show that AI-generated texts exhibit greater lexical diversity and syntactic complexity; however, they often exhibit structural uniformity, overuse of cohesive devices, and limited pragmatic depth, and should not replace professionally designed educational materials.

V. Smaglii, T. Korolova, Svitlana Yukhymets et al. · 0 citations
Open access Jul 2026

Texts Generated by Artificial Intelligence: Structure and Semantics

It was concluded that texts generated by artificial intelligence constitute a separate linguistic phenomenon with its own set of characteristics, which requires a special typology and a flexible, updatable analysis methodology.

L. Kravets, Viktória Stefuca, N. Libak et al. · 0 citations
Open access Aug 2026

Predicting Authorship Attribution in AI-Generated vs Human-Authored Texts: A Corpus-Based Study of Syntactic Complexity in Personal Statements

Generative AI complicates the use of personal statements as evidence of applicants’ individual voice. This corpus-based quantitative study examined whether syntactic-complexity measures distinguish human-authored from ChatGPT-generated personal statements for business and economics fellowship applications. The corpus comprised 50 publicly accessible human-authored texts and 50 texts generated by GPT-4o from a single prompt. Lu’s L2 Syntactic Complexity Analyzer yielded nine variables: word, sentence, and clause counts; dependent clauses per clause (DC/C) and T-unit (DC/T); complex T-units per T-unit (CT/T); coordinate phrases per clause (CP/C) and T-unit (CP/T); and T-units per sentence (T/S). Descriptive statistics, one-way MANOVA, follow-up ANOVAs, and discriminant function analysis were applied. The multivariate effect of text type was significant, Wilks’ Λ = .075, F(9, 90) = 123.51, p < .001, partial η² = .925; the discriminant function was also significant, χ²(9, N = 100) = 242.31, p < .001. Human texts were much longer (M = 915.80 vs. 150.68 words) and showed higher DC/C, DC/T, CT/T, and T/S values; CP/C and CP/T did not differ. Thus, text length and selected subordination and T-unit measures distinguished the corpora. Because the texts were not length-matched and the design used one model, one prompt, and no cross-validated classification, the findings are corpus-specific and cannot support a general-purpose AI-text detector.

M. Ayadi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.