Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this *wild* AI text comes...
Jenna Russell, Ben Glickenhaus, Katherine Thai et al.· 0 citations
This work evaluates Overthink on proprietary and open-source reasoning models across the FreshQA, SQuAD, and MuSR datasets, and shows that newer generations of RLMs, while showing a drastic increase in per-token cost, also exhibit up to a 2.3x increase in reasoning tokens, leaving them more vulnerable to Overthink atta...
Abhinav Kumar, Jaechul Roh, Ali Naseh et al.· arXiv.org· 92 citations· ⚡9
The results suggest that length generalization is a meaningful stress test for creative-writing models and a useful lens for distinguishing otherwise close models.
Initial human feedback reveals that AI-written novels contain interesting descriptions and concepts, but often fail in long-range coherence and prose quality, including conceptual repetition, distracting details, and weak dialogues.