Skip to content
#small language model Open access

Testing six LLM-generated sleep-EEG predictions on open polysomnography: a case study with a pre-frozen protocol and negative results

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research) · 2 references
Sleep and related disorders Sleep and Wakefulness Research

Abstract

Language models can write research hypotheses much faster than anyone can test them. Earlier studies have judged such hypotheses by expert ratings, benchmarks with known answers or laboratory experiments. We report a small, fully documented case from a different angle: what happens to a handful of model-generated predictions when they are turned into prespecified numerical tests on public data. We tested six operationalized sleep-EEG predictions selected from archived GLM-5.3 outputs, which supplied no bibliographic citations for those predictions. Before this study downloaded its data we froze a pre-analysis protocol in version control, with pass and refutation thresholds, controls and a label hierarchy; it was not deposited in a public registry, and the author had seen hypnograms and some EEG of the main dataset in an earlier, unrelated analysis (section 6). The main data were Sleep-EDF Expanded (Sleep Cassette, 78 healthy people, night 1); the Dreem DOD-H dataset (25 people, five independent scorers) was planned for scorer-agreement controls. Two hypotheses met their prespecified refutation condition. For the claim that spindle-band (12 to 16 Hz) bursts become more frequent in the last ten minutes of wake before the first N2 epoch, the median burst-rate ratio in 25 people was 0.71, below the refutation threshold of 1.2; the decision uses the point median, as prespecified, and the 95% interval (0.19 to 1.35) still includes 1.2, so this is a verdict of the protocol's rule rather than a firm disproof. For the claim that spindle amplitude grows with the time since the previous spindle, the median amplitude ratio in 77 people was 0.99 (95% CI 0.97 to 1.01) and the median Spearman rho was -0.04; because amplitude also decided which spindles our threshold detector counted, this applies to the detected spindles only. Three hypotheses were not decided because the sample minimum was not met after event-level qualification (17 and 12 people against 20; 1 record against 100); the model's fixed -75 µV amplitude rule, applied to the negative peak of the bipolar Fpz-Cz derivation, may have contributed to this attrition, which we did not test. One could not be tested, because no dataset we audited provides bilateral leg EMG in enough healthy sleepers. No hypothesis passed. Most of the effort went into turning each hypothesis into a decidable rule, and on standard open sleep data four of the six ended undecided rather than true or false. Several of the underlying mechanisms have published antecedents, which we compare with each prediction. This study contributes a documented case of operationalizing and evaluating six model-generated sleep predictions, including negative results, eligibility attrition and dataset limitations. The event detectors were checked only on synthetic signals, so the ability of the tests to detect a real effect was not established. The study does not estimate the accuracy of LLM-generated hypotheses or establish that the underlying physiological ideas are new. It is not a clinical study. Files. Kulma_2026_D1_Sleep_LLM_Hypotheses_v0.7.pdf is the report. The ZIP archive contains the report in Markdown, the pre-analysis protocol with its full history and deviations ledger, the analysis code with unit tests, result files, aggregation scripts, provenance of the generated hypotheses, every AI review round of the protocol, code and results, a manifest linking each claim to its supporting file, licences and data attribution, a short guide to reproduction (README.md) and SHA-256 checksums. Seventeen files are labelled public copies in which a few operational details were replaced; the manifest gives the hash of each original and of its copy. No recordings or hypnograms are redistributed; download scripts for both datasets are included. Some supporting notes are in Polish. Use of AI tools. The hypotheses were generated by GLM-5.3; the analysis code and drafts of the text were written with Claude Code (Anthropic) under the author's direction; OpenAI Codex and Google Antigravity were used for review and pre-release checks. These automated reviews do not replace review by a sleep electrophysiologist; no sleep electrophysiologist has read this report yet. AI tools are not authors; the author takes responsibility for all content (report, section 8). Preprint, not peer reviewed. Text and results: CC BY 4.0; raw model output: CC0 1.0; code: MIT. Datasets keep their own licences (see LICENSE.md).

View source

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.