Testing six LLM-generated sleep-EEG predictions on open polysomnography: a case study with a pre-frozen protocol and negative results
Language models can write research hypotheses much faster than anyone can test them. Earlier studies have judged such hypotheses by expert ratings, benchmarks with known answers or laboratory experiments. We report a small, fully documented case from a different angle: what happens to a handful of model-generated predi...