functional-standard-atlas: attenuation-corrected, territory-resolved benchmarking of variant effect predictors against saturation genome editing
Abstract
functional-standard-atlas: attenuation-corrected, territory-resolved benchmarking of variant effect predictors against saturation genome editing Software and frozen data underlying the manuscript "Measurement reliability bounds functional benchmarks and relocates where variant effect prediction fails" (N. Zhang). WHAT CHANGED IN THIS VERSIONThis release adds a cross-paper operating-point comparison (Supplementary Note S15, generated by the new lr_basis_check analysis). It reports the four splice and coding likelihood ratios quoted in Fig. 6B on two bases — at the scanned 95%-specificity threshold this study uses, and interpolated to exactly 95% specificity, the basis of the companion splice-region study — together with the full 69-cell comparison. No Fig. 6B value, evidence band, existing supplementary number, figure, or other reported result changes; the frozen atlas and all analyses are unchanged. CONTENTS- The frozen functional-standard atlas of 64,178 saturation genome editing variants across seven cancer-susceptibility genes (46,392 SNVs and 17,786 indels), uniformly processed on MANE-Select transcripts.- All nineteen predictor score columns.- Per-stratum reliability and attenuation-ceiling estimates.- The predictor-definition sweep and the reliability-correction simulation.- The MaveDB deposit survey.Environments are pinned exactly; raw and frozen data carry SHA-256 manifests; the pipeline is platform-independent and requires Python 3.12 or newer. DATA PROVENANCEThe saturation genome editing measurements are third-party public deposits, cited by accession and not redistributed here: MaveDB urn:mavedb:00000662-0-1 (BAP1), 00001250-a-2 (BARD1), 00000097-0-2 (BRCA1), 00001225-a-1 (BRCA2), 00001259-a-2 (PALB2), 00000673-0-1 (RAD51C) and 00000675-a-1 (VHL). RELATEDCompanion deposit (splice-region calibration benchmark), all versions: https://doi.org/10.5281/zenodo.22674887Manuscript preprint: [bioRxiv DOI — add after posting] Use of large language models: large language model assistance was used to write analysis code and in preparing the manuscript. No data were generated by the model. LICENSING — PLEASE READ BEFORE REUSEThe licence field of this record (MIT) covers the code only. The frozen assay matrix (data/frozen/frozen-matrix-v1.parquet, which holds no predictor column) is released under CC BY 4.0, compatible with all seven upstream MaveDB deposits. The score matrices and everything derived from them carry mixed terms, column by column, across nineteen third-party predictors. Eight columns are more restrictive than CC BY 4.0 (AlphaGenome, CADD, AlphaMissense, Nucleotide Transformer, SpliceAI, REVEL, PrimateAI, VEST4; AlphaMissense and Nucleotide Transformer are ShareAlike as well as NonCommercial), eight require attribution only, and for three (BayesDel, ClinPred, MetaRNN) no licence statement could be located, which must not be read as permission. Per-column terms are in results/predictor_resources_v1.tsv; the full statement is in LICENSE-DATA.