Omicau Multi-Omics Benchmark Suite
Abstract
Overview A prospectively frozen benchmark suite for leakage-safe multi-omic integration. It evaluates predictive performance, modality utility, null behavior, failure handling, and compute cost without making claims of clinical utility or causal biological inference. Included datasets Synthetic controls: paired null and planted-signal families for binary classification and continuous regression. DepMap/CCLE: transcriptomics, copy number, LC-MS metabolomics, and PRMT5 dependency across 644 cell lines. TCGA BRCA: transcriptomics, copy number, and RPPA protein abundance for ductal-versus-lobular classification across 783 tumors. TCGA LGG, KIRC, and UCEC: transcriptomics and copy number for IDH status, pathological stage, and histology endpoints across 507, 507, and 500 tumors, respectively. Design and controls All real cohorts were fixed after source and endpoint eligibility checks and before method performance was observed. The design uses shared group-aware partitions, training-only preprocessing, five outer folds repeated three times for real cohorts, 40 independent synthetic replicates, ten target permutations per real cohort, 5,000 paired group bootstraps, and Holm adjustment across the five primary real-dataset contrasts. Literature-anchored controls are evaluated independently of method ranking. Failed, unfavorable, discordant, and indeterminate outcomes remain reportable. Reproducibility The archive contains the frozen protocol, immutable source registry, download and validation code, group-aware partitions, fixed comparator implementations, statistical aggregation, schemas, environment pins, and fault-injection tests. Raw molecular matrices, participant-level data, local paths, and benchmark results are excluded. Deviations Deviation 1 - Aggregation target normalization. Final aggregation converts read-only NumPy memory-mapped target vectors to base NumPy arrays before metric and bootstrap validation. This preserves all values, ordering, dtypes, datasets, endpoints, partitions, methods, predictions, thresholds, statistical procedures, and frozen randomization streams. The correction has no scientific impact on the estimand. All definitive work units are regenerated under the corrected implementation identity; no prior run outputs are reused.