Jul 2026
ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation
Results show that detector performance measured on the conventional human-vs-LLM benchmark does not transfer to human-authored text revised by an LLM, even though the same detectors remain largely robust to LLM-only rewriting.
G. Perrone, S. Romano
· arXiv.org · 0 citations