variant-fm-benchmark: evidence strength of splice-region variant effect predictors against saturation genome editing functional assays
Abstract
variant-fm-benchmark 3.0.1: code, data and every output of the study "Evidence strength of splice-region variant effect predictors assessed against saturation genome editing data in seven cancer susceptibility genes" (N. Zhang), with one command that rebuilds every output from public sources and checks it against the archived files. WHAT THIS ISThe analysis set of 8,853 intron-side single-nucleotide variants within 50 nucleotides of the exon boundary in BRCA1, BRCA2, BARD1, PALB2, RAD51C, VHL and BAP1, labelled by each saturation genome editing deposit's own functional classification; the pipeline that applies the published SpliceAI cut points of the ClinGen splicing recommendation, fits predictor-specific PP3/BP4 evidence thresholds by interval calibration, tests each in genes left out of the fit, and applies them to two genes outside the set (DDX3X and TP53); every table, figure and supplementary table of the study; and the checksums and provenance records of every input. ONE COMMANDUnpack this archive, or clone https://github.com/oneone00-11/variant-fm-benchmark at tag v3.0.1-submission, and run bash reproduce.sh It needs Python 3.12 and network access. It builds the pinned environment (requirements-evidence.lock.txt), downloads and checks the two public inputs below, runs all 22 stages (about 25 minutes of computation) and compares every output with the archived file. It ends REPRODUCED when every output matches byte for byte, REPRODUCED NUMERICALLY when every printed table and figure does and other files differ only in the last digits of their numbers (relative 1e-9; numpy and scipy call the operating system's maths library on macOS, which differs between chips and releases), or with a list of what differs. On the machine the outputs were made on (macOS arm64, Python 3.12.13) it ends REPRODUCED, the build time recorded in the analysis-set manifest aside; on a GitHub-hosted M1 (macOS 14) it ends REPRODUCED NUMERICALLY, for this release's tag: https://github.com/oneone00-11/variant-fm-benchmark/actions/runs/36666472719. On x86_64 Linux every computed value agrees to 1e-9; figure files differ in their bytes where the Arial font is absent, and two supplementary tables print two AlphaGenome splice score thresholds, which lie within floating-point precision of an observed score, to fewer significant figures. bash reproduce.sh --check is a five-minute check without data download: every checksum the manifests and provenance records carry, then the test suite. THE TWO PUBLIC INPUTS1. The ClinVar GRCh38 VCF of 15 June 2026 (clinvar_20260615.vcf.gz, 192 MB), from NCBI (https://ftp.ncbi.nlm.nih.gov/pub/clinvar/vcf_GRCh38/archive_2.0/2026/). This record also holds NCBI's file, unchanged, as a mirror; the command uses it only when NCBI cannot serve the exact file, and checks the same sha256 whichever source served it.2. functional-standard-atlas release 2.5.0 (https://doi.org/10.5281/zenodo.22751081), whose score matrix the analysis set is built from. The command uses only a copy that matches the release file for file. NOT REBUILT BY THE COMMANDModel scores that need their own environments are archived with checksums and provenance records and read as inputs: the SpliceAI re-score at the recommendation's distance and its event records, the Pangolin event records, the DDX3X and TP53 model scores, and the AlphaGenome Atlas columns, which need an API key. LICENSING — PLEASE READ BEFORE REUSEThe licence field of this record (MIT) covers the code only. The data files carry the terms of their sources, column by column, as LICENSE-DATA in the archive sets out. The functional measurements carry the terms of their MaveDB deposits (CC0 1.0 or CC BY 4.0). Every column derived from AlphaGenome (alphagenome, alphagenome_v061, avi, avi_splice_sites, avi_splice_site_usage, avi_splice_junctions) is under the AlphaGenome Terms of Service: non-commercial research only; the outputs must not be used to train other machine-learning models, and the predictions must not be used for clinical decision-making. Other predictor score columns carry their tools' own terms, several of them non-commercial (SpliceAI's trained models CC BY-NC 4.0; CADD non-commercial; Nucleotide Transformer CC BY-NC-SA 4.0); Supplementary Table S3 of the study and LICENSE-DATA give the terms column by column. The ClinVar file is NCBI's, in the public domain. WHAT CHANGED SINCE 3.0.0Three tests that rebuild outputs and compare them with the archived files now hold to the numerical tier across machines: two rebuilds on one machine must be byte-identical, the printed S9 table must match byte for byte, and the rest must agree to a relative 1e-9. With them the test suite passes on x86_64 Linux (368 passed, 14 skipped). The ClinVar mirror URL in the code points at this record. The pipeline code and every output are those of 3.0.0. WHAT CHANGED SINCE 2.0.0Version 2.0.0 archived the calibration study on the frozen matrix. Version 3.0.0 adds the evidence-strength study and its one-command reproduction; the earlier studies' code and data are unchanged and keep their own entry points.