Results show that detector performance measured on the conventional human-vs-LLM benchmark does not transfer to human-authored text revised by an LLM, even though the same detectors remain largely robust to LLM-only rewriting.
Abstract
Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has shown that rewriting and paraphrasing can degrade detector performance, it remains unclear whether performance measured on this conventional benchmark predicts detector behavior when human-authored content is rewritten by an LLM. To address this gap, we introduce Authorship-Rewriting Benchmark (ARB), built from 1,800 human source texts (600 each from XSum, WritingPrompts, and OpenWebText) and four open-weight generators (Llama-3.2-3B, Qwen2.5-7B, Mistral-7B, Gemma-2-9B). Each source item yields four matched variants: human-written (HUMAN), direct LLM generation (Free-LLM), LLM-rewritten human text (H2L), and same-generator LLM-rewritten LLM text (LLM2L). We evaluated five detectors (FastDetectGPT, Binoculars-falcon-7b, RADAR, BERT-Defense, RoBERTa-Defense) at a strict 1%-false-positive operating point (TPR@1%FPR). FastDetectGPT and Binoculars-falcon-7b detected 91.2% and 93.5\% of direct LLM text, but only 30.8% and 15.1% of human text an LLM had rewritten, a drop of 60-78 percentage points. The same detectors retained 78.3% and 83.0% recall when LLM text was rewritten by the same model, a much smaller decline of 10-13 points. RADAR followed the same pattern (66.8% to 12.2%), while BERT-Defense and RoBERTa-Defense stayed below 3% recall across all regimes. These results show that detector performance measured on the conventional human-vs-LLM benchmark does not transfer to human-authored text revised by an LLM, even though the same detectors remain largely robust to LLM-only rewriting.
A paired-prompt benchmark for human-versus-machine detection across English text, Python code, and mixed text–code documents shows that reliable deployment requires cross-domain evaluation, mixed-content testing, and calibration beyond in-distribution accuracy.
The key idea is to smooth adjacent token scores to reduce their variability, while using an adaptive Lepski-type rule to select the bandwidth according to the local authorship structure, and the proposed method achieves favorable mean square error performance in estimating the underlying signal.
Yangjun Lu, Hongyi Zhou, F. Spill et al.· arXiv.org· 0 citations
Human evaluators struggle to distinguish original scientific abstracts from AI-generated text, as AI-produced formal language appears neat and convincing; prior studies report reviewers correctly identify only 68% of AI-generated abstracts while misclassifying 14% of human texts. This study presents an exploratory, generator-specific evaluation of mDeBERTa v3 using zero-shot Natural Language Inference (NLI) classification, applied to Indonesian scientific abstracts synthesized via IndoT5-base-paraphrase rather than AI-generated text in general. A balanced 2,274-abstract dataset paired human abstracts (SINTA 3 journals) with IndoT5-base-paraphrase outputs as the AI class. Mann-Whitney U analysis on seven linguistic features revealed significant differences (p < 0.001) across all. A critical anomaly emerged: AI texts showed higher sentence-length variation (SD = 15.44) than human texts (SD = 7.99), contradicting the assumption that AI text is more uniform, attributable to context-window exhaustion in IndoT5 producing semantic hallucinations when synthesizing dense abstracts. Testing three NLI scenarios showed a single instruction targeting this fluctuation achieved the highest Recall (76.52%) but with 790 false positives among 1,137 human abstracts, limiting accuracy to 53.52%; added complexity further degraded AI-class recall due to vocabulary overlap. A Random Forest classifier trained on the same features achieved 91.21% accuracy (F1 = 0.9130), substantially outperforming the zero-shot approach and confirming the anomaly as a strong, learnable signal. These results indicate zero-shot NLI can partially track a generator's mechanical artifacts through a single targeted instruction, but remains insufficiently precise to separate machine-error fluctuation from natural human variation, and is not recommended for standalone academic-integrity screening without further refinement
Aldo Syahputra, Aris Wahyu Murdiyanto, Ulfi Saidata Aesyi· Indonesian Journal of Data a...· 0 citations
Given the current trend to employ large language models (LLMs) in almost any imaginable context, LLM-generated text detection and authorship attribution have become a pressing issue. Prior work has primarily focused on surface-level linguistic features, an approach shown to be susceptible to paraphrasing and other obfuscation techniques. In this paper, we go beyond the linguistic surface, extracting and analysing reasoning structures in LLM-generated texts with the goal of capturing more complex signals of LLM authorship. We propose a graph neural network approach that leverages reasoning graphs extracted by an argument mining pipeline, demonstrating improved robustness and generalisation over a traditional Longformer baseline. Our approach outperforms the baseline by up to 27 percentage points under the obfuscation attacks such as paraphrasing and backtranslation, and 19 percentage points when evaluated on the texts generated by the unseen model versions, simulating real-world conditions in which new LLM versions are continuously released.
Zlata Kikteva, Artur Romazanov, Annette Hautli-Janisz et al.· arXiv.org· 0 citations
EVIL-Detect, a multi-signal ensemble framework with conflict-aware fusion for NLPCC 2026 Shared Task 6, improves robustness under strong out-of-distribution shifts, achieving a macro-F1 score of 0.8888 and ranking first in the official evaluation.
Hongrui Bao, Hangyu Rong, Zhuo Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.