Skip to content
#gene editing Open access

Does adjustment for tumor proliferation shift prognostic gene expression associations in a way that is predicted by each gene's correlation with proliferation? A confirmatory analysis across cancer types in The Cancer Genome Atlas

Sep 2026 · Open Science Framework

Abstract

1. What is already done, and what this document covers I state this first, because the value of a preregistration written after related work depends on being exact about what was already seen. Three analyses are already complete and are exploratory. None of them was preregistered, and none of them is covered by this document. A genome-wide sweep in TCGA-LIHC, in which the hazard ratio of each expressed gene was estimated with and without a proliferation score. The ratio of the two hazard ratios, which I call the shift, was regressed on the gene's Spearman correlation with the proliferation score. The slope was -0.278 with an R-squared of 0.785 in a model adjusted for age, sex, stage and grade, and -0.320 with an R-squared of 0.840 in a model adjusted for age, sex and stage. A replication of that sweep in GSE14520, an independent liver cohort, with a slope of -0.180. A pilot in two further cancer types, TCGA-KIRC and TCGA-LUAD, with slopes of -0.292 and -0.293 and R-squared values of 0.752 and 0.876. Those five slopes are the reason I expect what I expect. They are the discovery, and I will report them as such. This document preregisters a confirmatory analysis in the cancer types that have not yet been examined, so that the general claim rests on estimates whose analysis plan was fixed before they existed. TCGA-LIHC, TCGA-KIRC and TCGA-LUAD are excluded from the confirmatory set and will be reported separately as the exploratory set. No result from those three cohorts will be counted toward any decision rule below. 2. Hypotheses H1, primary. In each included cancer type, the slope of the log shift on the gene's Spearman correlation with the proliferation score is negative. H2, primary. The shift is a property of the gene rather than of the cohort, so that the shifts estimated in two different cancer types for the same gene are positively correlated. H3, secondary and descriptive. The number of genes that lose significance under Benjamini-Hochberg control when the proliferation score is added varies across cancer types. I make no directional prediction about its size, because the pilot showed that this varies from near total in TCGA-LIHC to modest in TCGA-KIRC. This hypothesis exists so that the variation is reported rather than discovered later and presented as a finding. 3. Data Expression: the TCGA HiSeqV2 matrices hosted by UCSC Xena, as log2(norm_count + 1), one matrix per cancer type. Clinical covariates: the corresponding TCGA clinical matrices hosted by UCSC Xena. Survival: overall survival and its time, from the pan-cancer Clinical Data Resource table hosted by UCSC Xena. No new human or animal data are generated. All data are public and de-identified, so ethical approval and informed consent are not required. 4. Which cancer types enter I do not name the cancer types in advance, because naming them invites selection. I fix the rule instead, and the script applies it to every TCGA cancer type for which UCSC Xena hosts a HiSeqV2 matrix. Every cancer type that the rule excludes will be listed with its reason. A cancer type is included when all of the following hold: A HiSeqV2 expression matrix and a clinical matrix are both retrievable. The clinical matrix carries an age field and a pathologic or clinical stage field. At least six of the eight proliferation score genes are present in the expression matrix. After restriction to primary tumors, one sample per patient, and complete data on age, stage and survival, at least 200 patients remain. At least 60 deaths are observed in that set. The thresholds of 200 patients and 60 deaths are chosen so that a Cox model with three or four covariates is not fitted on fewer than fifteen events per covariate. They are fixed here and will not be moved. 5. Analysis plan 5.1 The proliferation score The score is the mean of the gene-wise standardized expression of MKI67, TOP2A, CCNB1, PCNA, BUB1, CCNA2, AURKA and RRM2, computed within each cancer type, then standardized. This is the same score used in the exploratory work and is not re-derived. 5.2 The genes swept Within each cancer type, a gene is swept when its median and its interquartile range across the analysis set are both above zero. The eight genes of the proliferation score are excluded, because adjusting a gene for a score that contains it is circular. 5.3 The two models For each swept gene, two Cox proportional hazards models are fitted for overall survival. The first contains the standardized gene expression, age, sex and stage dichotomized as I to II against III to IV. The second adds the standardized proliferation score. Sex is dropped in a cancer type in which it does not vary. Stage is dropped in a cancer type in which the dichotomy does not vary, and that fact is reported. Grade is not used, because it is not recorded in every cancer type. I note that grade is not a substitute for proliferation: in the exploratory liver analysis, grade carried no prognostic information once age, sex and stage were accounted for. 5.4 The shift and the landscape The shift for a gene is the exponential of the difference between its coefficient in the second model and its coefficient in the first. The landscape is summarized by the ordinary least squares regression of the log shift on the gene's Spearman correlation with the proliferation score, within each cancer type. The reported quantities are the slope, its p value, and the R-squared. 5.5 Cross-cohort agreement For every pair of included cancer types, the Spearman correlation of the log shift is computed across the genes swept in both. The summary quantity is the median of those pairwise correlations. 5.6 Multiplicity Within each cancer type, the p values of the gene coefficients are adjusted by the Benjamini-Hochberg procedure across all swept genes, separately for the model without and the model with the proliferation score. 6. Decision rules These are fixed now and are the only basis on which I will describe the landscape as general. Hypothesis Confirmed if Not confirmed if H1 The slope is negative in at least 90 percent of included cancer types, and the median R-squared across them is at least 0.50 The slope is negative in fewer than 75 percent of included cancer types, or the median R-squared is below 0.30 H2 The median pairwise Spearman correlation of the log shift is at least 0.30 The median pairwise Spearman correlation is below 0.10 An outcome that falls between the two columns is reported as equivocal and described as such. I will not adopt a new threshold after seeing the result. If H1 is confirmed and H2 is not, the claim I will make is that the direction and size of the shift are reproducible at the level of the cancer type but not at the level of the individual gene, and the recommendation to readers will be stated accordingly. 7. What would falsify the underlying claim The claim is that omitting proliferation from a prognostic gene expression model biases the estimate in a direction and by an amount set by the gene's correlation with proliferation. It is falsified if the slope is not consistently negative across cancer types, or if it is so close to zero that the implied bias is smaller than the resampling noise of a single estimate. A slope that is negative but flat, say above -0.05, would mean the phenomenon is real but too small to change practice, and I would report it that way. 8. Analyses this document does not cover The exploratory sweeps in TCGA-LIHC, GSE14520, TCGA-KIRC and TCGA-LUAD, all of which are already run and will be labeled exploratory. The non-collapsibility analysis, which is already run in TCGA-LIHC and is exploratory. The survey of how often published prognostic gene studies in hepatocellular carcinoma adjust for proliferation. Its own design was fixed before any paper was read, and that design is reported with it, but it was not registered here and I will not describe it as preregistered. Any analysis of a specific gene, including CYB5R3. If I run anything beyond the plan above, it will be labeled post hoc in the manuscript, as the work already done is. 9. Deviations Any departure from this plan will be reported in the manuscript with the reason and the date, and the result under the original plan will be reported beside it. The analysis script writes this plan to a file on every run, so that an edit made after the fact is visible in the deposited output. 10. Timing and availability This document is registered before the confirmatory analysis is run. The analysis code and all result files will be deposited in the project's Zenodo record, and the record's digital object identifier will be cited in the manuscript.

View source

Similar papers

#computer vision Conference Aug 2008

A Preliminary Roadmap for Empirical Research on Agile Software Development

Some claim that especially in the field of agile software development the research lags years behind of the practice. In this paper, we characterize the status and main challenges for research on agile software development, and propose a preliminary roadmap, focusing on providing more empirical research, primarily on e...

Torgeir Dingsøyr, T. Dybå, P. Abrahamsson · 92 citations · ⚡7
#computer vision Book Open access Mar 2017

On the Unhappiness of Software Developers

The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 84 citations · ⚡6
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Conference Aug 2008

Scrum in a Multiproject Environment: An Ethnographically-Inspired Case Study on the Adoption Challenges

Agile methods continue to gain popularity. In particular, the Scrum method appears to be on the verge of becoming a de-facto standard in the industry, leading the so called Agile movement. While there are success stories and recommendations, there is little scientifically valid evidence of the challenges in the adoptio...

A. Marchenko, P. Abrahamsson · 59 citations · ⚡11

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.