Skip to content
#gene editing Open access

Measurement reliability bounds functional benchmarks and relocates where variant effect prediction fails

Sep 2026 · bioRxiv · 0 citations
Biology

Abstract

Background Variant effect predictors are increasingly benchmarked against multiplexed assays of variant effect (MAVEs) rather than clinical labels, which removes label circularity but introduces a new problem: a correlation against a measurement cannot exceed the measurement’s own reproducibility, and precision varies sharply across the territories compared. Results We scored nineteen predictors across sixteen strata of a frozen atlas of 64,178 saturation genome editing variants in seven cancer-susceptibility genes. From published replicate scores and standard errors we estimated each territory’s reliability ceiling, and showed by simulation that the correction reduces error above a ceiling of about 0.45 and amplifies it below. Ceilings vary more across territory than predictors do, and correcting for them redraws the map at the splice extremes. The collapse at canonical splice sites is largely a property of the assay: the median shortfall relative to coding narrows from 1.7-to 1.4-fold; this convergence survives dropping BARD1 or PALB2 but inverts when BRCA1 is dropped, so we report all three leave-one-gene-out folds rather than claim gene independence, and the frontier parity rests on one deposit. Genuine failure lies 11–50 bp into the intron, which the uncorrected map presents as modest. Across MaveDB, 2,452 of 2,803 score sets carry, at the upper bound, what a reliability estimate needs, though a conventional column-name search finds only a tenth; among 674 human deposits with a computable ceiling, 29.9–51.8% fall below 0.90. Scored as classification against the assays’ own functional calls in three genes, the same predictors separate damaging from tolerated better than their correlations suggest, though none reaches the strongest evidence band at the 95%-specificity operating point. Conclusions Territory-resolved benchmarks should report a per-stratum reliability estimate, or state that the assay permits none. It asks nothing of depositors and applies today, at the upper bound, to most (87%) of MaveDB.

Read PDF

Similar papers

#computer vision Conference Aug 2008

A Preliminary Roadmap for Empirical Research on Agile Software Development

Some claim that especially in the field of agile software development the research lags years behind of the practice. In this paper, we characterize the status and main challenges for research on agile software development, and propose a preliminary roadmap, focusing on providing more empirical research, primarily on e...

Torgeir Dingsøyr, T. Dybå, P. Abrahamsson · 92 citations · ⚡7
#computer vision Book Open access Mar 2017

On the Unhappiness of Software Developers

The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 84 citations · ⚡6
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Conference Aug 2008

Scrum in a Multiproject Environment: An Ethnographically-Inspired Case Study on the Adoption Challenges

Agile methods continue to gain popularity. In particular, the Scrum method appears to be on the verge of becoming a de-facto standard in the industry, leading the so called Agile movement. While there are success stories and recommendations, there is little scientifically valid evidence of the challenges in the adoptio...

A. Marchenko, P. Abrahamsson · 59 citations · ⚡11

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.