Skip to content

Agreement Between Large Language Models and Humans in Research Proposal Review - Data and Scripts

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Agreement Between Large Language Models and Humans in Research Proposal Review — Data and Code This repository contains the data and code required to reproduce the analyses, statistical results, and figures presented in the associated manuscript. Files are organized by function and described below. All research proposals are anonymized and labeled with non-identifying identifiers (A, B, C, …). Reviewer identities were never provided to the authors. To prevent inadvertent disclosure, all free-text review content from both human reviewers and large language models (LLMs) has been removed; only the numerical evaluation data required to reproduce the reported analyses are included. Version history Version 2 (revision, October 2026). 05_Data_Processing.ipynb was revised during peer review (notebook version 3). The data files are unchanged. The revised notebook: excludes the 1,866 self-exposed reviews of the two in-context exemplar proposals from all analyses (see Data provenance and known limitations); adds proposal-level bootstrap 95% confidence intervals for every model's ICC(2,1) and Spearman r_s, and paired model-minus-benchmark differences computed within the same resamples; adds a leave-one-out single-reviewer benchmark (ICC(2,1) and Spearman r_s) and class-level comparisons of reasoning-enabled and non-reasoning models; adds an analysis of the effect of exemplar self-exposure; replaces the combined ICC/Spearman figure with separate three-panel figures (manuscript Figs 1 and 2); writes the supporting tables S2–S5 as CSV files. Version 1. Archive accompanying the original submission. Data files Human_raw_scores.csv Individual numerical scores assigned by human reviewers, one row per reviewer × proposal × criterion. Used to compute panel-level summary statistics, the inter-reviewer reliability metrics, and the leave-one-out single-reviewer benchmark reported in the manuscript. Human.csv Proposal-level human reference scores (one row per proposal) used as the human benchmark against which LLM scores and rankings are compared. LLM_data_combined_clean_filtered.csv All numerical scores generated by the evaluated LLMs. Of the 52,301 evaluations generated, the 336 in which a model failed to return one or more required numerical scores were removed, leaving the 51,965 rows in this file. The file still contains the 1,866 self-exposed exemplar reviews; 05_Data_Processing.ipynb excludes them before analysis, so the analyzed dataset comprises 50,099 reviews (see Data provenance and known limitations). review_criteria.txt The evaluation criteria and rating scales. See the note under Data provenance regarding the 2023 vs. 2024 criterion naming. Data dictionary Human_raw_scores.csv Applicant: Anonymized proposal identifier (A, B, C, …). Cycle: Review cycle the proposal belongs to (2023 or 2024). Label: Scoring criterion: Intellectual merit, Potential for impact, Collaborative Potential, or Overall ranking. Reviewer_Seq: Reviewer index within a proposal (1, 2, 3, …). Identities are unknown; this index only links a single reviewer's ratings across criteria for the same proposal, in source-file order. It is not consistent across proposals (reviewer 1 for proposal A is not reviewer 1 for proposal B). Rating: Numerical score. Criterion ratings use a 1–5 scale; Overall ranking uses a 1–3 scale (3 = fund, 2 = fund with modifications, 1 = do not fund). Human.csv Name: Anonymized proposal identifier (matches Applicant above). Cycle: Review cycle (2023 or 2024). IM, Impact, Collab, Overall: Panel-mean scores for the four criteria. Score: Weighted composite panel score, computed as 0.4·IM + 0.3·Impact + 0.3·Collab, matching the weighting applied to the LLM composite scores. LLM_data_combined_clean_filtered.csv Name: Anonymized proposal identifier (matches Human.csv). Type: Input given to the model: Abstract or Full_Proposal. Model: Model identifier. For models with controllable reasoning depth, the tier is appended as a suffix (_low, _medium, _high). The portion before the first underscore is the root model. In the figures and S2–S4 Tables, results averaged across the three tiers are labeled "(R-avg)". Prompt: Prompting strategy. OneShot is zero-shot prompting; CoT is in-context exemplar (few-shot) prompting, in which two exemplar proposals and their human reviews were shown before the target proposal. The manuscript uses the terms "zero-shot" and "in-context exemplar"; the CoT label is historical, and no chain-of-thought instructions were given. Seed: Requested random seed. For models that did not support seed specification at the time of execution (the reasoning models listed in the Methods), the API ignored this value and it functions only as a replicate index; output is not reproducible from it for those models. Temp: Sampling temperature (0.1, 0.5, 0.9). IM, Impact, Collab: Criterion ratings (1–5 scale). Overall: Overall recommendation (1–3 scale). See limitation note on out-of-scale values. Score: Weighted composite, 0.4·IM + 0.3·Impact + 0.3·Collab. Category: Reasoning/architecture category of the model. Year: Evaluation wave in which the run was performed (proposals were re-evaluated as new model generations were released); this is not the proposal's submission cycle. Use Cycle in the human files for submission cycle. Data provenance and known limitations We document the following so that users can interpret the data accurately. Two review cycles, combined. The 28 proposals come from two internal seed grant cycles: 15 from 2023 and 13 from 2024 (Cycle column). For 2024 proposals, proposal-level means in Human.csv are the official institute panel means; for 2023 proposals they are computed from the individual ratings in Human_raw_scores.csv. Reviewers per proposal. Panels ranged from 3 to 6 reviewers. Because one review is missing at the individual level for six 2024 proposals (next paragraph), Human_raw_scores.csv contains 2 to 6 reviews per proposal, 122 in total (mean 4.36). Reviewer identities were never provided; the human inter-reviewer reliability is therefore estimated with a one-way random-effects model (ICC(1,1)), which is the appropriate model when each proposal is rated by a different, unidentified set of reviewers. Six 2024 reviews not present at the individual level. For six 2024 proposals (G, R, T, W, X, Z), one reviewer's scores were submitted without written comments and are not included in Human_raw_scores.csv. For these proposals, Human.csv carries the official institute panel means, so the panel mean in Human.csv and the mean recomputed from Human_raw_scores.csv differ slightly. The reproducibility check in 06_Human_data.ipynb confirms exact agreement for all proposals with complete individual records and reports the expected small differences for these six. In-context exemplar proposals and self-exposed reviews. Under the in-context exemplar condition (Prompt = CoT), proposals O and B, the highest- and lowest-scoring proposals of the 2023 cycle, were shown to the model with their human reviews before each target proposal. These two proposals were also scored as targets under the same condition, so in 1,866 reviews the model had been shown the target proposal's own human review immediately beforehand. These rows are retained in LLM_data_combined_clean_filtered.csv so that their effect can be reproduced. Block 1 of 05_Data_Processing.ipynb excludes them from all analyses (EXEMPLAR_PROPOSALS = ["O", "B"]), and Block 9 quantifies the effect of the exposure and of the exclusion (S5 Table). Out-of-scale LLM Overall ratings. A small number of responses (627 of the 50,099 analyzed reviews, 1.25%, almost entirely from gpt-3.5-turbo) rated the overall recommendation on a 1–5 scale rather than the requested 1–3 scale, in a format the parser could not distinguish. The analysis code masks values outside the valid range before computing any Overall-based result; the composite Score does not use Overall and is unaffected. Excluded LLM evaluations. The 336 evaluations in which a model failed to return one or more required numerical scores were removed before analysis and are not included in the file provided here. Code notebooks Two groups of notebooks are provided: (i) the LLM evaluation pipeline and (ii) statistical analysis and figure generation. LLM evaluation pipeline (Notebooks 1–4) Documentation of the methodology used to generate the LLM evaluations. These use synthetic examples and contain no confidential data, API credentials, or real proposal text. 01_pipeline_overview.ipynb — architecture, configuration, criteria, output format, evaluation matrix. 02_prompting_strategies.ipynb — zero-shot (OneShot) and in-context exemplar (CoT) prompting; exemplar selection; text vs. vision input. 03_response_parsing.ipynb — regex extraction of ratings and comments; error handling; decimal ratings. 04_example_evaluation.ipynb — end-to-end workflow on synthetic data. Statistical analysis and figures (Notebooks 5–6) 05_Data_Processing.ipynb (version 3, revised September 2026) — reads LLM_data_combined_clean_filtered.csv, Human.csv, and Human_raw_scores.csv. Run the blocks in order: Block 1: Load data; exclude self-exposed exemplar reviews; mask out-of-scale Overall values; Type II ANOVA and partial η² (Table 1) Block 2: Empirical expected absolute differences (EAD) (S1 Table) Block 3: Average scores across prompting strategy, seed, and temperature (reasoning tiers kept separate) Block 4: Merge model scores with human panel means Block 5: ICC(2,1) and Spearman r_s point estimates; R-avg aggregation; display-model ordering Block 6: Bootstrap 95% CIs (proposal resampling); human ICC(1,1) and ICC(1,k); paired differences vs. ICC(1,1) (S2 Table) Block 7: Leave-one-out single-reviewer benchmark; paired differences; class-level comparisons (S3 Table, S4 Table) Block 8: Agreement figures: ICC(2,1) and Spearman r_s with CIs and paired diffe

View source

Similar papers

#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Conference Open access Dec 2013

Affordable and Energy-Efficient Cloud Computing Clusters: The Bolzano Raspberry Pi Cloud Cluster Experiment

The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.

P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al. · 110 citations · ⚡7
#computer vision Book Open access Mar 2017

On the Unhappiness of Software Developers

The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 84 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.