COLD READ: The Anonymity Half-Life Is a Property of the Reader
Abstract
How many words can you write before a language model can infer who you are? This work went looking for that number and found that the question is malformed, which is itself the finding. 72 authors from the Blog Authorship Corpus, balanced across three age bands and both genders, were shown to local language models in growing slices of their own text (25 to 1600 words). At every step the model was forced to commit to gender, age band, and star sign. Star sign is the negative control: it is labelled in the corpus and is not inferable from prose. It never cleared its bar at any step in any of the three models, which is the load-bearing check that makes the rest of the measurement trustworthy. Main result: the same 72 authors, the same words and the same prompt produce wildly different exposure curves depending on the model doing the reading. llama3.1:8b clears a coin flip on gender (Wilson lower bound above 0.50) at 50 words and reaches 90.3% by 1600; qwen2.5:7b-instruct needs roughly 800 words; mistral:7b clears only at 1600. All three are in the same 7-8B size class, so family alone produces a thirty-two-fold difference on identical text. No claim of the form "you are anonymous for N words" is meaningful without naming the model, and N falls as models improve. Secondary results: age band clears earlier than gender in all three models. The first two plateaued near 60%, which suggested a ceiling set by the text; the third reached 70.8% and was still climbing, so that ceiling claim is withdrawn, and the readers do not share one ranking (the slowest gender reader is tied fastest on age band). Exposure accrues smoothly rather than snapping on at a threshold. Size matters within a family too: llama3.2 at 3B, run on the identical protocol, never clears chance on gender within 1600 words, where llama3.1:8b clears at 50; age band clears at 100 words against 25. The star-sign control held again. Three of its 504 calls never finished (deterministic runaway generation at temperature 0) and are reported as missing. The two sizes are also two releases (3.2 and 3.1), so the contrast is size plus one release step. A finding that did not replicate, reported as such: the first model read below chance at short lengths, suggesting that short samples surface stereotype matching rather than uncertainty. The second model was above chance from the first step and the third dipped only at the first slice; no interval at that slice excludes chance, so the claim is not supported, and the write-up says so explicitly rather than quietly dropping it. Contamination was tested directly and ruled out rather than argued away, by scoring model continuations of a verbatim prefix against the author's real next words versus a different author's, with the probe itself verified to be functioning before its null result was accepted. Position against prior work: the profiling task is not new. Argamon, Koppel, Pennebaker and Schler (Automatically profiling the author of an anonymous text, Communications of the ACM, 2009) established that age and gender are recoverable from ordinary prose, on this same corpus. Nor is the length axis new: Eder (Does size matter? Authorship attribution, small samples, big problem, Digital Scholarship in the Humanities, 2015) showed attribution accuracy depends on sample length and collapses below a minimum, sweeping length against a fixed classifier. In the LLM era, Staab, Vero, Balunovic and Vechev (arXiv:2310.07298) measured attribute inference at near-human accuracy, and Lermen, Paleka, Swanson, Aerni, Carlini and Tramer (arXiv:2602.16800) demonstrated large-scale profile linkage; neither sweeps input size. The contribution here is the interaction those literatures hold fixed on one side or the other: length swept across three different readers of one size class, where the threshold moves thirty-two-fold on identical text, with a labelled negative control. Repository contains the sampling and analysis code, the raw JSONL results for all four readers, the contamination probe, and a consent-gated two-seat demonstration application.