Skip to content
#gene editing Open access

Omitting tumor proliferation from prognostic gene expression models: preregistered analysis code, decision rules and results

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)
Gene expression and cancer classification

Abstract

R analysis code, preregistered decision rules, result files, run logs and figures for a study of what omitting tumor proliferation does to prognostic gene expression estimates. For every expressed gene in a cohort, two Cox models are fitted for overall survival, one adjusted for age, sex and anatomic extent and one that adds a standardized eight-gene proliferation score. The ratio of the two hazard ratios is the shift. Within a cohort, the log shift is regressed on the gene's Spearman correlation with the score, and the slope of that line is the cohort's landscape slope, b. L is the log hazard ratio of the proliferation score in that same cohort's model. This version adds the primary test of the second registration. The plan doi:10.17605/OSF.IO/N8DW6, registered on 2026-09-20, predicted that a cohort's landscape slope is set by the prognostic strength of proliferation in that cohort. Version 2.4.0 stated that no analysis covered by that registration was in the repository. That is no longer true, and the README has been corrected. The test used twenty cohorts that no earlier stage had examined: four cancer types of The Cancer Genome Atlas re-screened at the thresholds the registration lowered, and sixteen series of the Gene Expression Omnibus assembled under rules written down before the search was run. Together they hold 4,986 patients, 1,681 deaths, fourteen cancer types, twelve expression platforms and 459,827 gene fits. All three registered hypotheses were confirmed. - H1, coefficient of L negative with an interval excluding zero: -0.8419 (95 percent interval -0.8998 to -0.7840). CONFIRMED.- H2, intercept with an interval including zero: -0.0090 (-0.0303 to 0.0123). CONFIRMED.- H3, sign concordance at least 80 percent: 20 of 20. CONFIRMED. The fit explains 0.9811 of the variance of b across cohorts. Set B alone gives -0.8831 (-0.9542 to -0.8119) and Set A alone -0.7985 (-1.0171 to -0.5799). The attenuation-corrected estimate is -1.0485 at a reliability ratio of 0.803. Seven of the twenty cohorts have a negative L, so the intercept is read inside the observed range rather than extrapolated to it, and its interval excludes 0.0643, the magnitude of non-collapsibility measured independently in the discovery cohort. CU05_combined_decision.csv records the verdict. The first registration, doi:10.17605/OSF.IO/X5DCF, remains refuted and is still reported as refuted. Nothing in this version changes that verdict. What is new in the tree. PREREG2_NEXT/ holds the Stage N, Q, R1, R1b, R2a and R2b scripts, the field and annotation rules they share, the figure script and a working record of each step. Seventy-two result files were added, under the prefixes CN, CQ, CR, CS, CT, CU and CV. The figure is deposited in five formats. How the cohorts were chosen is itself deposited. The registration did not name accession numbers, so the search rules were written to a result file before the first query was sent. All 182 hits are recorded, all 161 screened series are listed with the reason for each decision, and the adjudication of every candidate, 287 rows, names the time, status, age and extent columns each cohort would use before any expression matrix was read. One series states in its own description that its status field codes the event as zero, so reading it by the usual convention would have inverted survival in that cohort. That is why this step is a recorded human decision rather than an inference. Corrections in this version. The comment header of analysis/36_approximation.R attributed an objection to a referee; this manuscript has never been peer reviewed and the objection was raised in the author's own pre-submission check, so the header now says so, with no change to any number. A working record gave the cancer-type count as eleven, having counted one cohort set only; the deposited result file carries fifteen disease labels, fourteen if colon and colorectum are treated as one organ, and the correction is annotated in place. The deposit title now matches the manuscript. Superseded script versions are kept beside their replacements as .v1, .v2 and .draft_20261007. They produced no deposited output. They are included because four of this round's corrections were to the author's own code rather than to the data, and the uncorrected versions are the evidence for that. Every exploratory and post hoc analysis is labeled as such in the code and in the output. Every script writes its own decision rules to a result file on each run, so a rule edited after the fact is visible in the deposited output rather than only in the code history. All primary data are public. Expression, clinical annotation and survival for the cohorts of The Cancer Genome Atlas come from UCSC Xena; the series of the Gene Expression Omnibus are listed by accession number in the deposited result files; gene sets come from MSigDB. No new human or animal data were generated and no identifiable data are included. License MIT. Cite the version DOI alongside the article.

View source

Similar papers

#computer vision Conference Aug 2008

A Preliminary Roadmap for Empirical Research on Agile Software Development

Some claim that especially in the field of agile software development the research lags years behind of the practice. In this paper, we characterize the status and main challenges for research on agile software development, and propose a preliminary roadmap, focusing on providing more empirical research, primarily on e...

Torgeir Dingsøyr, T. Dybå, P. Abrahamsson · 92 citations · ⚡7
#computer vision Book Open access Mar 2017

On the Unhappiness of Software Developers

The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 84 citations · ⚡6
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Conference Aug 2008

Scrum in a Multiproject Environment: An Ethnographically-Inspired Case Study on the Adoption Challenges

Agile methods continue to gain popularity. In particular, the Scrum method appears to be on the verge of becoming a de-facto standard in the industry, leading the so called Agile movement. While there are success stories and recommendations, there is little scientifically valid evidence of the challenges in the adoptio...

A. Marchenko, P. Abrahamsson · 59 citations · ⚡11

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.