eQTL weight-source choice reshapes TWAS gene candidacy: a two-axis dual-source audit with disease-agnostic calibration and a genome-wide benchmark (code & data)
Abstract
Code and processed data release accompanying the manuscript “eQTL weight-source choice reshapes TWAS gene candidacy: a two-axis dual-source audit with disease-agnostic calibration and a genome-wide benchmark” (BMC Genomics submission). v2.2.0 (2026-09-11) — numerically stable ACAT-O, fixed-threshold and architecture-unselected controls: (1) All GTEx multi-tissue ACAT-O p-values were recomputed with the numerically stable cotangent form T = Σ cot(pπ) / p = arctan(1/T)/π; the archived standard-form implementation returned a single clipped value of 2.11×10-15 for the 14 most extreme gene–phenotype pairs, and those superseded values are no longer reported (data/processed/gtex_acat_o_results.csv regenerated; Additional file 1: Table S2 regenerated). (2) Fixed-threshold enrichment reanalysis removing the stratum-size dependence of the Benjamini–Hochberg procedure (data/hk_reselect_20260830/data/d3_results.txt, manuscript Table S9). (3) Genome-wide, architecture-unselected random control: 600 genes drawn from the 11,820 genome-wide genes carrying a GTEx v8 Whole_Blood MASHR model, plus a 30-gene control and a 30-gene bootstrap null (d3_gw_sample_raw.csv, d3_gw_acat.csv, d3_results.txt; manuscript Table S10). (4) HRT-restricted architecture-unselected second control: the original HRT Atlas v1.0 source file data/hk_reselect_20260830/data/Human_Mouse_Common.csv is archived byte-faithfully; running rules R1–R5 of 01_select_hk_genes.py on it reproduces the recorded filter chain exactly (1129→1112→1013→1007→767), and relaxing R5 to Whole_Blood-only gives the 818-gene second-control pool (d3b_* files; manuscript Table S11). The 30-gene HRT-restricted control enriches at 65.3% (49/75) against a null median of 61.9%, corroborating the genome-wide control. (5) RPS16 leave-one-out sensitivity analysis. The archived HRT source file is required to reproduce the housekeeping control selection (Methods 1.8) and both architecture-unselected controls. [2026-09-14 update] [2026-09-14 update] Package refreshed to match the finalized manuscript (concept DOI 10.5281/zenodo.21238202). Figures are numbered Fig1-Fig8 plus FigS1-FigS6 exactly as in the manuscript (the earlier Figure2 was actually Fig. S1, which had shifted every later figure number by one, and the files have been renumbered accordingly). Fig. 4, 5, 6, 7 and 8 were regenerated: the earlier Fig. 5 read a stale intermediate file, the earlier Fig. 7 was a five-study merge that did not correspond to the manuscript, the earlier Fig. 8 carried a hard-coded correlation, and Fig. 6 is now the three-panel version; every plotted value is now computed from the harmonized source tables at run time, and the previous scripts that contained hard-coded numbers were removed. The table CSVs were regenerated verbatim from the finalized manuscript, replacing an earlier pipeline generation. A provenance README was added under figures/. No manuscript .docx or supplementary body text is included in this archive. [2026-09-15 update] [2026-09-15 update] Corrected Fig. 1. The in-figure pointer to the reporting checklist now reads "Additional file 1: Table S11"; it previously carried the superseded label "Table S7" from before the checklist was split into Tables S11 and S13. Fig1.png and Fig1.pdf have been re-exported (the PDF remains vector text), and the Fig. S1 caption in figures/FigS1_caption.txt was made identical to the caption in Additional file 1. No data, code, table or analytical result was changed in this release. [2026-09-20 update] [2026-09-20 update] Complete figure set, current figure scripts, and two data-layer corrections. (1) figures/ now ships the full manuscript figure set — Fig1-Fig8 plus FigS1-FigS6 (previously only FigS1) — byte-identical to the finalized submission files; all Fig1-Fig8 were re-exported after a format pass (page width ≤170 mm, minimum font size 6 pt, non-Type3 fonts, RGB PNG) and Fig. 1 gained the "diagnostic partition" qualifier. (2) figure_scripts/ now holds the officialZ pipeline — the only script set that reproduces the shipped figures: every path goes through a new paths_config.py (override with TWAS_REPO / TWAS_DATA_Z / FIG_OUT_MAIN / FIG_OUT_SUPP / AF1_DOCX / FIG_RESULTS), and the two json that previously existed only in the authors' session directories are now included: m15_positive_control.json (the input behind Fig. S6) and fig4_bootstrap_officialZ.json (the deterministic B = 5,000 bootstrap product behind Fig. 4b). The five scripts of the 2026-09-14 generation read the pre-correction data layer and reproduce retired values (RNH1 DR Z 13.318 vs 2.3064; TUBB 48.516 vs 11.8932; CKAP4 4.7549 vs 0.9874); they are retained only as a provenance note (README_数据来源_legacy_20260914.md) and must not be used to reproduce this work. (3) tables/ now carries Table 1a and 1b (the manuscript's Table 1 was split) plus Table 2, 3a, 3b, 3c and 4 exported verbatim from the current manuscript; the captions file is regenerated accordingly. (4) Two files of the processed-data layer were regenerated to match Additional file 1 exactly, having been one generation stale: gtex_official_Z.csv (FDR q columns, 217 of 222 rows) and crosscohort_TableS4_official.csv (pooled Z 1.51; Q 0.13 and I² 0% for the √N_e-weighted row). The figures are unaffected by (4). No manuscript or supplementary .docx is included in this archive. [2026-09-21 update] Correction of one heterogeneity statistic in the RNH1 cross-population row, so that this archive agrees cell by cell with the submitted manuscript and with the archived estimation code. (1) tables/Table4_integrated_evidence_assessment.csv: the reported I² of the RNH1 cross-population-replication row changes from 20.9% to 20.6%. 20.9% is not reproducible from the archived inputs of that row (FinnGen R13 DR with eQTLGen weights, Z = +2.3091; UK Biobank GCST90043640 with eQTLGen weights, Z = +0.7225). The archived DerSimonian–Laird implementation behind the row (Fig7_rnh1_crosscohort.py::dersimonian_laird, unit variances v = [1, 1], so C = 1 and tau^2 = max(0, Q − 1)) returns Q = 1.2587 and I² = 20.55%, i.e. 20.6% at one decimal, and is consistent with the Q = 1.26 carried in the same row. The same figure is corrected in Additional file 1 (Table S4a, row 1), in the main text (Results, the Table 4 note, and Limitation 4) and in the Chinese translation; the GitHub repository was aligned in the same pass (README.md, audit_notes/, and data/processed_officialZ/crosscohort_TableS4_official.csv, which is the parsed source of this value). The retired 20.9% required Q = 1.2642 (UK Biobank Z ≈ 0.7190), which corresponds to no archived input. tau is unaffected: both Q = 1.2587 and Q = 1.2642 round to tau = 0.51. (2) No other file in this archive changed: the remaining 72 files are byte-identical to version 2.7.0. [2026-09-21 update 2] tables/ re-exported from the current finalized manuscript, so that the per-table CSV snapshots in this archive match the submitted tables cell for cell. Three tables had drifted because the manuscript's tables were restructured after the previous package build (2026-09-20 18:07): (1) Table1a_disease_agnostic_calibration.csv — the fourth column is now the four-element "Screen / Source / Draw / Tissue" description of each control layer, replacing the earlier free prose; (2) Table3c_gene_subset_sensitivity.csv — row labels shortened and reordered, and the anchor-set row now states explicitly that it adds four non-testbed proteins and omits NCL and PDIA6; (3) Table4_integrated_evidence_assessment.csv — the table was reduced from seven columns to five: the "Mahalanobis Confirmation" and "Cross-Population Replication" columns were dropped in the manuscript and their content folded into the "Overall Verdict" and "Margin provenance" columns. Table_captions_and_notes.txt was regenerated against the same manuscript (11 blocks covering all seven tables and their notes, in place of the previous 13; the two former standalone "Note." blocks are now merged into the Table 2 and Table 3b/3c notes, matching the manuscript after its house-style pass, and the extraction is now anchored on the notes' own opening words rather than on a stale key list). Table 1b, 2, 3a, 3b and every figure, figure script and simulation file are byte-identical to version 2.8.0.