Agreement Between Large Language Models and Humans in Research Proposal Review - Data and Scripts
Abstract
Agreement Between Large Language Models and Humans in Research Proposal Review — Data and Code This repository contains the data and code required to reproduce the analyses, statistical results, and figures presented in the associated manuscript. Files are organized by function and described below. All research proposals are anonymized and labeled with non-identifying identifiers (A, B, C, …). Reviewer identities were never provided to the authors. To prevent inadvertent disclosure, all free-text review content from both human reviewers and large language models (LLMs) has been removed; only the numerical evaluation data required to reproduce the reported analyses are included. Version history Version 2 (revision, October 2026). 05_Data_Processing.ipynb was revised during peer review (notebook version 3). The data files are unchanged. The revised notebook: excludes the 1,866 self-exposed reviews of the two in-context exemplar proposals from all analyses (see Data provenance and known limitations); adds proposal-level bootstrap 95% confidence intervals for every model's ICC(2,1) and Spearman r_s, and paired model-minus-benchmark differences computed within the same resamples; adds a leave-one-out single-reviewer benchmark (ICC(2,1) and Spearman r_s) and class-level comparisons of reasoning-enabled and non-reasoning models; adds an analysis of the effect of exemplar self-exposure; replaces the combined ICC/Spearman figure with separate three-panel figures (manuscript Figs 1 and 2); writes the supporting tables S2–S5 as CSV files. Version 1. Archive accompanying the original submission. Data files Human_raw_scores.csv Individual numerical scores assigned by human reviewers, one row per reviewer × proposal × criterion. Used to compute panel-level summary statistics, the inter-reviewer reliability metrics, and the leave-one-out single-reviewer benchmark reported in the manuscript. Human.csv Proposal-level human reference scores (one row per proposal) used as the human benchmark against which LLM scores and rankings are compared. LLM_data_combined_clean_filtered.csv All numerical scores generated by the evaluated LLMs. Of the 52,301 evaluations generated, the 336 in which a model failed to return one or more required numerical scores were removed, leaving the 51,965 rows in this file. The file still contains the 1,866 self-exposed exemplar reviews; 05_Data_Processing.ipynb excludes them before analysis, so the analyzed dataset comprises 50,099 reviews (see Data provenance and known limitations). review_criteria.txt The evaluation criteria and rating scales. See the note under Data provenance regarding the 2023 vs. 2024 criterion naming. Data dictionary Human_raw_scores.csv Applicant: Anonymized proposal identifier (A, B, C, …). Cycle: Review cycle the proposal belongs to (2023 or 2024). Label: Scoring criterion: Intellectual merit, Potential for impact, Collaborative Potential, or Overall ranking. Reviewer_Seq: Reviewer index within a proposal (1, 2, 3, …). Identities are unknown; this index only links a single reviewer's ratings across criteria for the same proposal, in source-file order. It is not consistent across proposals (reviewer 1 for proposal A is not reviewer 1 for proposal B). Rating: Numerical score. Criterion ratings use a 1–5 scale; Overall ranking uses a 1–3 scale (3 = fund, 2 = fund with modifications, 1 = do not fund). Human.csv Name: Anonymized proposal identifier (matches Applicant above). Cycle: Review cycle (2023 or 2024). IM, Impact, Collab, Overall: Panel-mean scores for the four criteria. Score: Weighted composite panel score, computed as 0.4·IM + 0.3·Impact + 0.3·Collab, matching the weighting applied to the LLM composite scores. LLM_data_combined_clean_filtered.csv Name: Anonymized proposal identifier (matches Human.csv). Type: Input given to the model: Abstract or Full_Proposal. Model: Model identifier. For models with controllable reasoning depth, the tier is appended as a suffix (_low, _medium, _high). The portion before the first underscore is the root model. In the figures and S2–S4 Tables, results averaged across the three tiers are labeled "(R-avg)". Prompt: Prompting strategy. OneShot is zero-shot prompting; CoT is in-context exemplar (few-shot) prompting, in which two exemplar proposals and their human reviews were shown before the target proposal. The manuscript uses the terms "zero-shot" and "in-context exemplar"; the CoT label is historical, and no chain-of-thought instructions were given. Seed: Requested random seed. For models that did not support seed specification at the time of execution (the reasoning models listed in the Methods), the API ignored this value and it functions only as a replicate index; output is not reproducible from it for those models. Temp: Sampling temperature (0.1, 0.5, 0.9). IM, Impact, Collab: Criterion ratings (1–5 scale). Overall: Overall recommendation (1–3 scale). See limitation note on out-of-scale values. Score: Weighted composite, 0.4·IM + 0.3·Impact + 0.3·Collab. Category: Reasoning/architecture category of the model. Year: Evaluation wave in which the run was performed (proposals were re-evaluated as new model generations were released); this is not the proposal's submission cycle. Use Cycle in the human files for submission cycle. Data provenance and known limitations We document the following so that users can interpret the data accurately. Two review cycles, combined. The 28 proposals come from two internal seed grant cycles: 15 from 2023 and 13 from 2024 (Cycle column). For 2024 proposals, proposal-level means in Human.csv are the official institute panel means; for 2023 proposals they are computed from the individual ratings in Human_raw_scores.csv. Reviewers per proposal. Panels ranged from 3 to 6 reviewers. Because one review is missing at the individual level for six 2024 proposals (next paragraph), Human_raw_scores.csv contains 2 to 6 reviews per proposal, 122 in total (mean 4.36). Reviewer identities were never provided; the human inter-reviewer reliability is therefore estimated with a one-way random-effects model (ICC(1,1)), which is the appropriate model when each proposal is rated by a different, unidentified set of reviewers. Six 2024 reviews not present at the individual level. For six 2024 proposals (G, R, T, W, X, Z), one reviewer's scores were submitted without written comments and are not included in Human_raw_scores.csv. For these proposals, Human.csv carries the official institute panel means, so the panel mean in Human.csv and the mean recomputed from Human_raw_scores.csv differ slightly. The reproducibility check in 06_Human_data.ipynb confirms exact agreement for all proposals with complete individual records and reports the expected small differences for these six. In-context exemplar proposals and self-exposed reviews. Under the in-context exemplar condition (Prompt = CoT), proposals O and B, the highest- and lowest-scoring proposals of the 2023 cycle, were shown to the model with their human reviews before each target proposal. These two proposals were also scored as targets under the same condition, so in 1,866 reviews the model had been shown the target proposal's own human review immediately beforehand. These rows are retained in LLM_data_combined_clean_filtered.csv so that their effect can be reproduced. Block 1 of 05_Data_Processing.ipynb excludes them from all analyses (EXEMPLAR_PROPOSALS = ["O", "B"]), and Block 9 quantifies the effect of the exposure and of the exclusion (S5 Table). Out-of-scale LLM Overall ratings. A small number of responses (627 of the 50,099 analyzed reviews, 1.25%, almost entirely from gpt-3.5-turbo) rated the overall recommendation on a 1–5 scale rather than the requested 1–3 scale, in a format the parser could not distinguish. The analysis code masks values outside the valid range before computing any Overall-based result; the composite Score does not use Overall and is unaffected. Excluded LLM evaluations. The 336 evaluations in which a model failed to return one or more required numerical scores were removed before analysis and are not included in the file provided here. Code notebooks Two groups of notebooks are provided: (i) the LLM evaluation pipeline and (ii) statistical analysis and figure generation. LLM evaluation pipeline (Notebooks 1–4) Documentation of the methodology used to generate the LLM evaluations. These use synthetic examples and contain no confidential data, API credentials, or real proposal text. 01_pipeline_overview.ipynb — architecture, configuration, criteria, output format, evaluation matrix. 02_prompting_strategies.ipynb — zero-shot (OneShot) and in-context exemplar (CoT) prompting; exemplar selection; text vs. vision input. 03_response_parsing.ipynb — regex extraction of ratings and comments; error handling; decimal ratings. 04_example_evaluation.ipynb — end-to-end workflow on synthetic data. Statistical analysis and figures (Notebooks 5–6) 05_Data_Processing.ipynb (version 3, revised September 2026) — reads LLM_data_combined_clean_filtered.csv, Human.csv, and Human_raw_scores.csv. Run the blocks in order: Block 1: Load data; exclude self-exposed exemplar reviews; mask out-of-scale Overall values; Type II ANOVA and partial η² (Table 1) Block 2: Empirical expected absolute differences (EAD) (S1 Table) Block 3: Average scores across prompting strategy, seed, and temperature (reasoning tiers kept separate) Block 4: Merge model scores with human panel means Block 5: ICC(2,1) and Spearman r_s point estimates; R-avg aggregation; display-model ordering Block 6: Bootstrap 95% CIs (proposal resampling); human ICC(1,1) and ICC(1,k); paired differences vs. ICC(1,1) (S2 Table) Block 7: Leave-one-out single-reviewer benchmark; paired differences; class-level comparisons (S3 Table, S4 Table) Block 8: Agreement figures: ICC(2,1) and Spearman r_s with CIs and paired diffe