Omitting tumor proliferation from prognostic gene expression models: preregistered analysis code, decision rules and results
Abstract
R analysis code, preregistered decision rules, result files, run logs and figures for a study of what omitting tumor proliferation does to prognostic gene expression estimates. For every expressed gene in a cohort, two Cox models are fitted for overall survival, one adjusted for age, sex and anatomic extent and one that adds a standardized eight-gene proliferation score. The ratio of the two hazard ratios is the shift. Within a cohort, the log shift is regressed on the gene's Spearman correlation with the score, and the slope of that line is the cohort's landscape slope, b. L is the log hazard ratio of the proliferation score in that same cohort's model. This version adds the primary test of the second registration. The plan doi:10.17605/OSF.IO/N8DW6, registered on 2026-09-20, predicted that a cohort's landscape slope is set by the prognostic strength of proliferation in that cohort. Version 2.4.0 stated that no analysis covered by that registration was in the repository. That is no longer true, and the README has been corrected. The test used twenty cohorts that no earlier stage had examined: four cancer types of The Cancer Genome Atlas re-screened at the thresholds the registration lowered, and sixteen series of the Gene Expression Omnibus assembled under rules written down before the search was run. Together they hold 4,986 patients, 1,681 deaths, fourteen cancer types, twelve expression platforms and 459,827 gene fits. All three registered hypotheses were confirmed. - H1, coefficient of L negative with an interval excluding zero: -0.8419 (95 percent interval -0.8998 to -0.7840). CONFIRMED.- H2, intercept with an interval including zero: -0.0090 (-0.0303 to 0.0123). CONFIRMED.- H3, sign concordance at least 80 percent: 20 of 20. CONFIRMED. The fit explains 0.9811 of the variance of b across cohorts. Set B alone gives -0.8831 (-0.9542 to -0.8119) and Set A alone -0.7985 (-1.0171 to -0.5799). The attenuation-corrected estimate is -1.0485 at a reliability ratio of 0.803. Seven of the twenty cohorts have a negative L, so the intercept is read inside the observed range rather than extrapolated to it, and its interval excludes 0.0643, the magnitude of non-collapsibility measured independently in the discovery cohort. CU05_combined_decision.csv records the verdict. The first registration, doi:10.17605/OSF.IO/X5DCF, remains refuted and is still reported as refuted. Nothing in this version changes that verdict. What is new in the tree. PREREG2_NEXT/ holds the Stage N, Q, R1, R1b, R2a and R2b scripts, the field and annotation rules they share, the figure script and a working record of each step. Seventy-two result files were added, under the prefixes CN, CQ, CR, CS, CT, CU and CV. The figure is deposited in five formats. How the cohorts were chosen is itself deposited. The registration did not name accession numbers, so the search rules were written to a result file before the first query was sent. All 182 hits are recorded, all 161 screened series are listed with the reason for each decision, and the adjudication of every candidate, 287 rows, names the time, status, age and extent columns each cohort would use before any expression matrix was read. One series states in its own description that its status field codes the event as zero, so reading it by the usual convention would have inverted survival in that cohort. That is why this step is a recorded human decision rather than an inference. Corrections in this version. The comment header of analysis/36_approximation.R attributed an objection to a referee; this manuscript has never been peer reviewed and the objection was raised in the author's own pre-submission check, so the header now says so, with no change to any number. A working record gave the cancer-type count as eleven, having counted one cohort set only; the deposited result file carries fifteen disease labels, fourteen if colon and colorectum are treated as one organ, and the correction is annotated in place. The deposit title now matches the manuscript. Superseded script versions are kept beside their replacements as .v1, .v2 and .draft_20261007. They produced no deposited output. They are included because four of this round's corrections were to the author's own code rather than to the data, and the uncorrected versions are the evidence for that. Every exploratory and post hoc analysis is labeled as such in the code and in the output. Every script writes its own decision rules to a result file on each run, so a rule edited after the fact is visible in the deposited output rather than only in the code history. All primary data are public. Expression, clinical annotation and survival for the cohorts of The Cancer Genome Atlas come from UCSC Xena; the series of the Gene Expression Omnibus are listed by accession number in the deposited result files; gene sets come from MSigDB. No new human or animal data were generated and no identifiable data are included. License MIT. Cite the version DOI alongside the article.