Skip to content
#small language model Open access

COLD READ: The Anonymity Half-Life Is a Property of the Reader

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

How many words can you write before a language model can infer who you are? This work went looking for that number and found that the question is malformed, which is itself the finding. 72 authors from the Blog Authorship Corpus, balanced across three age bands and both genders, were shown to local language models in growing slices of their own text (25 to 1600 words). At every step the model was forced to commit to gender, age band, and star sign. Star sign is the negative control: it is labelled in the corpus and is not inferable from prose. It never cleared its bar at any step in any of the three models, which is the load-bearing check that makes the rest of the measurement trustworthy. Main result: the same 72 authors, the same words and the same prompt produce wildly different exposure curves depending on the model doing the reading. llama3.1:8b clears a coin flip on gender (Wilson lower bound above 0.50) at 50 words and reaches 90.3% by 1600; qwen2.5:7b-instruct needs roughly 800 words; mistral:7b clears only at 1600. All three are in the same 7-8B size class, so family alone produces a thirty-two-fold difference on identical text. No claim of the form "you are anonymous for N words" is meaningful without naming the model, and N falls as models improve. Secondary results: age band clears earlier than gender in all three models. The first two plateaued near 60%, which suggested a ceiling set by the text; the third reached 70.8% and was still climbing, so that ceiling claim is withdrawn, and the readers do not share one ranking (the slowest gender reader is tied fastest on age band). Exposure accrues smoothly rather than snapping on at a threshold. Size matters within a family too: llama3.2 at 3B, run on the identical protocol, never clears chance on gender within 1600 words, where llama3.1:8b clears at 50; age band clears at 100 words against 25. The star-sign control held again. Three of its 504 calls never finished (deterministic runaway generation at temperature 0) and are reported as missing. The two sizes are also two releases (3.2 and 3.1), so the contrast is size plus one release step. A finding that did not replicate, reported as such: the first model read below chance at short lengths, suggesting that short samples surface stereotype matching rather than uncertainty. The second model was above chance from the first step and the third dipped only at the first slice; no interval at that slice excludes chance, so the claim is not supported, and the write-up says so explicitly rather than quietly dropping it. Contamination was tested directly and ruled out rather than argued away, by scoring model continuations of a verbatim prefix against the author's real next words versus a different author's, with the probe itself verified to be functioning before its null result was accepted. Position against prior work: the profiling task is not new. Argamon, Koppel, Pennebaker and Schler (Automatically profiling the author of an anonymous text, Communications of the ACM, 2009) established that age and gender are recoverable from ordinary prose, on this same corpus. Nor is the length axis new: Eder (Does size matter? Authorship attribution, small samples, big problem, Digital Scholarship in the Humanities, 2015) showed attribution accuracy depends on sample length and collapses below a minimum, sweeping length against a fixed classifier. In the LLM era, Staab, Vero, Balunovic and Vechev (arXiv:2310.07298) measured attribute inference at near-human accuracy, and Lermen, Paleka, Swanson, Aerni, Carlini and Tramer (arXiv:2602.16800) demonstrated large-scale profile linkage; neither sweeps input size. The contribution here is the interaction those literatures hold fixed on one side or the other: length swept across three different readers of one size class, where the threshold moves thirty-two-fold on identical text, with a labelled negative control. Repository contains the sampling and analysis code, the raw JSONL results for all four readers, the contamination probe, and a consent-gated two-seat demonstration application.

View source

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.