Skip to content
#small language model Open access

Problem Choice Without a Referee: An Automated Novelty Check Certifies an Empty Search, a Rediscovery Benchmark Credits the Hypothesis Humans Found Next, and Every Located Referee of Worth Independent of Proposer and Field Needs the Objective Given

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0. Pipelines that claim to automate scientific discovery gate their search on a judgement that a proposed idea or problem is novel and worth pursuing. The best-controlled evidence on that judgement points two ways at once: in a blind study involving more than one hundred NLP researchers, expert reviewers rated language-model research ideas as more novel than expert-written ones, and when forty-three researchers executed randomly assigned ideas from the same study, the model ideas lost that advantage and fell further than expert ideas on every metric, novelty included. This paper asks what, if anything, can referee a proposed research problem before the field has worked on it. It distinguishes four kinds of referee: an absence check (was a match retrieved?), a judgement before execution (does a reader rate it novel or promising?), uptake (did humans later publish it?), and outcome (did executing it work against an objective fixed in advance?). Assembling published measurements for each, it argues that the absence check fails mainly at retrieval, with comparison a smaller and less-measured source of error, and is benchmarkable without labels; that model and expert judges disagree about which questions are non-obvious, and human judges have documented biases of their own, including lower merit scores for highly novel proposals; that uptake benchmarks supply the research background and score the rediscovered hypothesis, so by construction they credit only what humans went on to publish, as several of their authors concede; and that, of the referees located, the only one that scores worth and is independent of both the proposer and the field's present and future judgement is the outcome referee, which requires the objective to be supplied. Choosing the objective is what problem discovery means, so at that step the only referees of worth on offer are the field's own judgement, now or later. That conclusion is partly definitional, and the observation that current systems take the research question as given has been made before; the contribution is the decomposition, the measurements that locate each referee's failure, the finding that the absence check can be audited without labels, and a refutation condition that a domain-general criterion of problem worth could meet. The claim is scoped to research-idea and research-question generation; it does not say that such systems cannot discover anything. The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every arXiv citation was machine-verified against its live arXiv Atom API record, and every other citation against its Crossref record (with PubMed, OpenAlex or Semantic Scholar used to read abstracts Crossref does not carry), during drafting; title and author list were checked against the record returned. The full texts (arXiv PDF renderings) of the load-bearing sources were read for the passages and numbers attributed to them, including Si, Yang and Hashimoto (2024), Si, Hashimoto and Yang (2025), Gupta and Pruthi (2025), Beel et al. (2025), Lu et al. (2024), Yamada et al. (2025), Sinhahajari et al. (2026), Liu and Zhai (2026), Wen et al. (2025), Sourati and Evans (2023), Krenn et al. (2023), Luo et al. (2025), Yang et al. (MOOSE-Chem), Kumar et al. (2025), Wang et al. (RND) and Bianchi et al. (2025). Every quantitative claim is taken from the abstract, main text or a table of the source credited with it; where published numbers are summed, the text says so. No experiment was run and no number in this paper was measured by its author. Table 1 and Figure 1 re-present published values, each named with its source. The four-referee decomposition, Table 2, Algorithm 1 and the proposed tests in Section 12 are conceptual synthesis by the author, not empirical results, and are presented as such.

View source

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.