Skip to content
Review

Detecting AI-Generated Text: Mechanisms, Robustness, and the Limits of Reliable Detection

Jul 2026 · International Journal of Research Publication and Reviews · 0 citations

TL;DR

It is concluded that AI-text detection, in its current form, cannot serve as a sole, dispositive basis for academic-integrity decisions, and a set of institutional and technical recommendations are proposed that better address the underlying problem than detector accuracy alone can.

Abstract

Large language models (LLMs) have become capable of producing human-like prose, institutions ranging from universities to publishers have adopted automated AI-text detectors — most visibly Turnitin's AI writing indicator, GPTZero, and similar tools — as a control against misrepresenting machine-generated content as human work. This paper examines how these detectors work, the empirical and theoretical evidence on their reliability, and the documented techniques that allow AI-generated text to evade them. Drawing on the peer-reviewed and preprint literature, we describe three detector families (zero-shot statistical detectors, trained classifiers, and watermarking schemes), summarize adversarial results showing that paraphrasing attacks can collapse watermark detection true-positive rates from above 99% to under 10%, and review evidence that current detectors produce systematically higher false positive rates for non native English writers — in one widely cited study, 61.3% of TOEFL essays were misclassified as AI generated versus near zero misclassification of native-speaker essays. We conclude that AI-text detection, in its current form, cannot serve as a sole, dispositive basis for academic-integrity decisions, and we propose a set of institutional and technical recommendations — process-based evidence, disclosed AI-use policies, watermarking at the model level, and human-adjudicated review — that better address the underlying problem than detector accuracy alone can.

View source

Similar papers

Review Open access Aug 2026

Detection of AI-Obfuscated, AI-Refined, and Humanized AI-Generated Text: A Systematic Review

This systematic literature review synthesizes peer-reviewed and high-quality studies published between 2023 and 2026 on AI-obfuscated, AI-refined, and humanized text and suggests that AI-text detection should be treated as one supportive signal rather than a stand-alone judgment, particularly in high-stakes academic or professional contexts.

Batyr Sharimbayev, S. Kadyrov · 0 citations
Open access Aug 2026

Misclassification of text in Ai detection: A serious limitation of Ai detectors and Its threats to human-based scholarly writing

Dear Editor, The rapid integration of AI into academic writing has necessitated tools for detecting AI-generated content such as iThenticate ZeroGPT, Turnitin, Phrasly AI, Open AI text Classifier, Writer, Copy leaks to differentiate between human and machine-generated text.1 However, current AI detection tools suffer from significant misclassification rates, generating false positives that wrongly accuse human authors and false negatives that allow AI-generated text to evade detection.2 This unreliability unduly impacts non-native English speakers, who often utilise AI tools for paraphrasing and grammar correction to ensure their work meets academic standards; however, these legitimate linguistic refinements are frequently misconstrued by detection software as evidence of AI-generated content.3 This problem is heightened by the fact that many AI detection tools, primarily trained on English corpora, struggle to accurately assess texts with diverse linguistic structures and stylistic conventions, thereby increasing the risk of false positives for non-English speaking scholars.2 Such false positives, where human-written manuscripts are incorrectly labelled as AI-generated , severely threaten scholarly psychological safety by fostering distrust and creating an environment of anxiety and unfair allegations. This systemic issue fundamentally challenges academic integrity, undermining the credibility of authors and the foundation of human-based scholarly writing.4,5 This editorial highlights the urgent need for human centric approach to AI detection, advocating for strategies that prioritise human-AI collaboration over sole reliance on fallible automated systems. A fundamental flaw of current AI detectors lies in their documented inconsistency and unreliability. Studies consistently demonstrate high rates of both false positives, where human-written text is erroneously identified as AI-generated, and false negatives, where AI-generated content evades detection.6,7 For instance, literature indicated that both free and commercially available AI detection tools can incorrectly classify human-written content as AI-generated with rates ranging from 43.3% to 83.3%.7,8 The ethical concerns arise directly because of incorrectly labelling human-written manuscripts as AI-generated and vice versa. Such errors lead to unfair allegations, rejection of genuine work, unwarranted accusations of academic misconduct, and reputational damage for authors.9,10 When authors must modify their writing or use "humaniser" tools to avoid false detection, it paradoxically increases AI involvement and further obscures human-machine authorship boundaries.8 Furthermore, the ease with which AI-generated text can be altered allows it to bypass current detection methods, turning the process into a counterproductive 'cat-and-mouse' game.1 This eventually weakens the very goal of identifying AI misuse while simultaneously unjustly burdening diligent human scholars.4 ---Continue

Noureen Durrani, Yusra Nasir, Sobia Ali · 0 citations
#artificial intelligence Preprint Aug 2026

IndicDetect: Evaluating Cross-Lingual LLM-Generated Text Detection for Hindi, Telugu, and Tamil

IndicDetect provides standard data splits, an evaluation protocol, and baselines to establish a robust, language-aware foundation for AI-generated text detection in Indic scripts, and finds substantial robustness failures.

Bhaskar Ganesh Devalla, Junchao Wu, Nilesh Dokuparthi et al. · 0 citations
Review Sep 2026

The adversarial game between detection and evasion: A survey of anti-detection techniques for machine-generated texts.

With the explosive growth of large language models (LLMs), research on machine-generated text detection (MGTD) has also proliferated. Alongside these developments, a wide range of attack algorithms targeting MGTD systems have emerged. While previous studies have surveyed detection techniques, few have examined the dynamic interplay between attack and defense. Following PRISMA 2020, this paper systematically synthesizes 27 studies of attacks against MGTD and the available evidence on corresponding defenses. We categorize existing research into four major types of evasion strategies: watermark attacks, paraphrasing attacks, prompt-based attacks, and adversarial-text attacks, and summarize the available defense evidence. Furthermore, to better understand the practical implications of these methods, we compile the reported performance results of attack and defense techniques across different detectors. Finally, we highlight the current challenges in this area and outline potential future research directions. A companion repository containing the categorized literature, paper links, and available code, data, and project repositories is provided at https://github.com/AIGC1999/A-Survey-of-Anti-Detection-Techniques-for-Machine-Generated-Texts.

Unknown authors · 0 citations
Review Sep 2026

Probability is not proof: Why AI detection differs from plagiarism detection in academic misconduct cases

This conceptual synthesis examines why artificial intelligence (AI) text detection cannot be treated as equivalent to plagiarism detection in cases of academic misconduct. Plagiarism detection tools operate forensically: they identify specific passages, match them to verifiable sources, and produce evidence that both instructors and students can examine and rebut. AI detectors operate statistically: they estimate the probability that a text resembles AI-generated writing, without identifying a source, act, or comparator. Drawing on independent evaluations showing no tool exceeds 80 per cent accuracy,1 false-positive rates above 61 per cent for non-native English writers in a Test of English as a Foreign Language (TOEFL)-essay testing condition,2 and journalistic survey evidence reporting disproportionate false-accusation experiences among Black students,3 the paper applies three lenses. First, evidentiary: AI scores are not falsifiable and suffer base-rate problems that make predictive value unknowable in real classrooms. Secondly, procedural: under Goss v. Lopez (1975) and Mathews v. Eldridge (1976), students are entitled to notice and a meaningful opportunity to respond, a safeguard undermined when the evidence is an opaque probability. Thirdly, fairness: applying Rawlsian justice, rational agents behind a veil of ignorance would reject a system whose errors fall most heavily on non-native speakers and on groups that survey evidence suggests may face disproportionate risk of accusation. Recent reporting on one federal preliminary ruling suggests discipline is more defensible when institutions rely on corroborating evidence, not scores alone. The paper concludes that detection outputs should never serve as standalone proof and recommends process-based assessment, bias audits, and transparent policies. This article is also included in The Business & Management Collection which can be accessed at https://hstalks.com/business/.

Unknown authors · 0 citations
Preprint Aug 2026

Why AI Detection Fails for Academic Integrity

Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled study of published English abstracts (four domains; 2013 to 2015 vs. 2023 to 2025), we quantify this policy failure under proxy human/AI labels at tau=0.50. Light"refine abstract only"edits, a proxy for guideline-compliant AI assistance, are flagged at 38 to 80%. Unmodified 2023 to 2025 originals are flagged at 9 to 15%, with non-STEM rates far above STEM (p<0.001); elevated scores track long-token and Academic Word List density, not authorship intent alone. After Undetectable AI humanization, evasion is near-total: fewer than 4% of AI-labeled rewrites remain flagged (post-humanization detection rate<4%; FNR>96%). Honest AI-editing results in a higher sanction risk than humanizer-assisted evasion. Therefore, detector scores should not serve as standalone misconduct evidence.

Jonathan A. Karr, Grigorii Khvatskii, T. Hua et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.