Skip to content
#protein folding Open access

Omicau Multi-Omics Benchmark Suite

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Overview A prospectively frozen benchmark suite for leakage-safe multi-omic integration. It tests predictive performance, modality utility, model-capacity effects, missingness handling, null behavior, exploratory external transport, failure handling, and compute cost without making claims of clinical utility or causal biological inference. Included datasets Synthetic controls: paired null and planted-signal families with operative structured missingness for binary classification and continuous regression. DepMap/CCLE: transcriptomics, copy number, LC-MS metabolomics, and PRMT5 dependency across 644 cell lines. TCGA BRCA: transcriptomics, copy number, and RPPA protein abundance for ductal-versus-lobular classification across 783 tumors. TCGA LGG, KIRC, and UCEC: transcriptomics and copy number for IDH status, pathological stage, and histology endpoints across 507, 507, and 500 tumors, respectively. CPTAC UCEC exploratory holdout: transcriptomics and copy number for endometrioid-versus-serous transport assessment across 95 patient-disjoint tumors. Methods and controls Nine fixed methods compare Omicau with an unmasked architecture-matched ablation, matched early and single-modality neural controls, early and single-modality linear controls, weighted late fusion, a missingness-only diagnostic, and calibrated latent partial least squares. A TCGA-UCEC complete-training-feature sensitivity tests outcome-associated technical missingness. Internal cohorts use shared group-aware partitions, training-only preprocessing, five outer folds repeated three times, 5,000 paired group bootstraps, paired DeLong tests for AUROC, 4,999 paired squared-error sign flips for R-squared, and Holm adjustment across five primary matched-capacity contrasts. Effect sizes and intervals are the primary evidence. Because repeated out-of-fold predictions share training sets, internal p-values are conditional on the frozen prediction vectors and are not unconditional population-generalization tests. Ten target permutations per cohort are coarse catastrophic-leakage diagnostics, not formal empirical tail-probability tests. TCGA and CPTAC expression scales are harmonized by a source-declared, target-blind transformation with pooled and matched-feature numerical-domain gates. The CPTAC endpoint is exploratory because it was exercised during predeposit development smoke. Literature-anchored controls remain independent of method ranking. Failed, unfavorable, discordant, non-estimable, and indeterminate outcomes remain reportable. Reproducibility The archive contains the frozen protocol, immutable source registry, download and validation code, internal and external partitions, fixed comparator implementations, statistical aggregation, schemas, environment pins, and fault-injection tests. Raw molecular matrices, participant-level data, local paths, and benchmark results are excluded. Deviations Deviation 1 - Aggregation target normalization. Final aggregation converts read-only NumPy memory-mapped target vectors to base NumPy arrays before metric and bootstrap validation while preserving scientific values and frozen randomization streams. Deviation 2 - Comparator and ablation expansion. Matched neural, unmasked, missingness-only, complete-feature, weighted late-fusion, and partial least-squares controls separate fusion value from model capacity, technical missingness, and integration strategy. All settings are fixed before definitive execution. Deviation 3 - Independent external evaluation. A CPTAC UCEC holdout adds 95 patient-disjoint assessment cases. TCGA-UCEC supplies all training and model-selection rows; shared transcriptomic and copy-number features are aligned by unique Entrez identifiers and expression scales are harmonized without using CPTAC outcomes. Deviation 4 - Primary contrast realignment. The primary contrast is Omicau minus the matched early neural control. Prior linear comparisons remain fully reported as contextual estimates and are not substituted for the matched-capacity test. Deviation 5 - External development exposure. The CPTAC endpoint was exercised during predeposit development smoke. External estimates are designated exploratory and are not treated as untouched confirmatory validation. Deviation 6 - External expression-scale correction. Predeposit development smoke exposed incompatible TCGA and CPTAC expression domains. Linear TCGA RSEM values now receive log2(x+1) after negative values are marked missing; CPTAC retains its source-declared log2 scale. Target-blind pooled and matched-feature gates validate compatibility. The correction precedes definitive execution, and earlier smoke outputs are not reused. Deviation 7 - Synthetic missingness application correction. Predeposit audit showed that registered synthetic missingness masks were not reaching model matrices. The masks now alter every synthetic method input exactly as registered. The correction precedes definitive execution, and earlier smoke outputs are not reused.

View source

Similar papers

#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Book Open access Jul 2015

Understanding the affect of developers: theoretical background and guidelines for psychoempirical software engineering

This paper highlights the challenges to conduct proper affect-related studies with psychology, provides a comprehensive literature review in affect theory, and proposes guidelines for conducting psychoempirical software engineering.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 56 citations · ⚡4
#machine learning Open access May 2017

What Influences the Speed of Prototyping? An Empirical Investigation of Twenty Software Startups

This study conducts a multiple case study on twenty European software startups and proposes a prototype-centric learning model in early stage software startups, and identifies factors that occur as barriers but also facilitators for prototyping in earlystage software startups.

Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson · 44 citations · ⚡5
#protein folding Open access Sep 2026

Programmable design of functional proteins from natural language

Pinal, a 16-billion-parameter foundation model that produces protein candidates from natural-language functional descriptions, supports natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or sequence constraints.

Fengyuan Dai, Shiyang You, Yudian Zhu et al. · 31 citations · ⚡3

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Google DeepMind Blog Nov 25, 2025

AlphaFold: Five years of impact

Explore how AlphaFold has accelerated science and fueled a global wave of biological discovery.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.