Skip to content
Open access

Leveraging multiplicity in biologically informed neural networks to uncover disease heterogeneity

Jul 2026 · bioRxiv · 0 citations · 32 references
Biology

TL;DR

Analysis of the space of BINN explanations can turn multiplicity into a tool for studying complex disease mechanisms, and across 100 replicate BINNs for type 2 diabetes, the authors find distinct solution clusters prioritizing either inflammatory or hepatic-metabolic pathways.

Abstract

Biologically inspired neural networks (BINNs) embed pathway, ontology, or protein-interaction structure directly into neural networks, promising interpretable disease prediction where hidden nodes map to named biological entities. Yet BINNs have been hard to train at biobank scale, and the reliability of their interpretations remains largely untested. Here we present a fast BINN implementation trained on UK Biobank genotype and plasma proteomics data from about 500,000 individuals across six common diseases. BINNs achieve competitive predictive performance, but we uncover two major limits to their interpretability. First, attribution scores are strongly biased by graph topology, because node degree and layer position influence the scores. Normalization reduces this bias but can weaken enrichment for known disease genes. Second, BINNs show substantial predictive multiplicity, that is, independently trained models with identical architecture and data reach similarly accurate solutions while prioritizing different genes and pathways. Although this multiplicity makes single-model explanations unstable, the range of interpretations can itself reveal disease biology. Across 100 replicate BINNs for type 2 diabetes, we find distinct solution clusters prioritizing either inflammatory or hepatic-metabolic pathways, mirroring known disease heterogeneity. Thus, analyzing the space of BINN explanations can turn multiplicity into a tool for studying complex disease mechanisms.

Read PDF

Similar papers

Open access Jul 2026

An interpretable omnigenic neural network architecture for the human genome

Genetic prediction of complex phenotypes typically relies on additive linear models, which scale well but cannot capture non-additive effects or deeply integrate molecular and clinical data. Domain-specific neural networks have driven advances in images, text, and other modalities, but genome-scale neural networks remain challenging because genotypes are sparse and high-dimensional, effective sample sizes are limited, and generic architectures lack interpretability. Here, we introduce the omnigenic neural network, a biologically structured architecture inspired by the omnigenic model of complex traits. The model learns hierarchical representations of biological processes, accommodates multimodal inputs, supports transfer learning, and enables multitask prediction. Models trained in the UK Biobank and evaluated in the All of Us cohort for ischemic heart disease, type 2 diabetes, and schizophrenia outperformed published PGS Catalog and PRS-CSx scores. A multitask model trained across 36 cardiovascular endpoints further outperformed corresponding single-phenotype models and baselines. The architecture provides systems-level interpretability by quantifying the contributions of biological processes, which were consistent with established disease mechanisms. It also captures non-linear interactions between variants. Analysis of these interactions using Integrated Hessians revealed patterns concordant with previously reported epistatic associations. Together, these findings establish the omnigenic neural network as a flexible framework for interpretable, multimodal, and multitask genomic prediction.

J. Upmeier zu Belzen, L. Arnoldt, N. Hollmann et al. · 0 citations
Open access Aug 2026

Biology‐informed neural networks learn nonlinear representations from omics data to improve genomic prediction and biological discovery

SUMMARY Traditional genotype‐to‐phenotype models depend heavily on direct mappings that achieve only modest accuracy, forcing breeders to conduct large, costly field trials to maintain or marginally improve genetic gain. Models that incorporate intermediate molecular phenotypes can achieve higher predictive fit, but remain impractical since such data are unavailable at deployment or design time. Biology‐informed neural networks (BINNs) overcome this limitation by encoding pathway‐level inductive biases and leveraging multi‐omics data only during training, while using genotype data alone during inference. Here, we extend BINNs for genomic prediction and selection in crops by integrating thousands of single‐nucleotide polymorphisms with multi‐omics measurements and prior biological knowledge. By directly embedding omics‐derived priors, BINN outperforms conventional models in low‐data (n < p) regimes and enables sensitivity analyses that expose biologically meaningful traits. Applied to maize gene expression and multi‐environment field trial data, BINN improves rank correlation accuracy within and across most subpopulations under sparse data conditions and nonlinearly identifies genes that GWAS/transcriptome‐wide association studies may fail to uncover. With complete domain knowledge for a synthetic metabolomics benchmark, BINN substantially reduces prediction error relative to conventional neural nets and correctly identifies the most important nonlinear pathway. Importantly, both cases show that highly sensitive BINN latent variables correlate with the experimental quantities they represent, despite not being trained on them. This suggests that BINNs learn biologically relevant representations, nonlinear or linear, from genotype to phenotype. Together, BINNs establish a framework for improved genomic prediction accuracy and biological discovery that can guide genomic selection, candidate gene selection, pathway enrichment, and gene‐editing prioritization.

Katiana Kontolati, R. J. Gladstone, Ian Davis et al. · 0 citations
Open access Jul 2026

FloREN: Decoding Immune Regulatory Networks through Interpretable Graph Transformer Patient Representations

A Framework for Learning Over REgulatory-Embedding Networks (FloREN), a supervised and interpretable sample representation method that enables improved sample stratification and biomarker discovery and supports downstream analyses that found specific immune network mechanisms in immune-mediated inflammatory diseases (IMIDs).

Iñigo Clemente‐Larramendi, S. Hillion, D. Cornec et al. · 0 citations
Open access Aug 2026

Biochemically Constrained Multi‐Omics Integration Reveals Protein–Metabolite Dependencies Across Diseases

ABSTRACT Integrating proteomic and metabolomic data is essential for understanding complex diseases, yet current approaches that rely primarily on statistical associations often overlook the structured biochemical relationships between molecular entities and suffer from discriminative instability in small clinical cohorts. Here, we present ProMetNet, a biochemically constrained framework that incorporates pathway‐derived connectivity from the Reactome database into neural network architecture. By encoding protein–metabolite relationships based on reaction topology, ProMetNet models structured cross‐omics dependencies rather than relying solely on statistical correlations, reducing spurious associations while preserving global molecular context and improving robustness in data‐limited settings. Across four heterogeneous disease cohorts, including Alzheimer's disease, type 2 diabetes, COVID‐19, and glioblastoma, ProMetNet consistently outperforms evaluated multi‐omics integration methods, including MOGONET, P‐NET, PEARL, and MOINER, maintaining high discriminative performance under substantial data downsampling. In addition to classification accuracy, the framework prioritizes biologically plausible protein–metabolite dependencies that are not captured by conventional differential or correlation‐based analyses. Importantly, pathway‐level signals identified by ProMetNet demonstrate consistent discriminative performance in independent large‐scale population data from the UK Biobank (N = 47,507), supporting their robustness and generalizability. Together, these results establish ProMetNet as a biologically grounded and interpretable framework for multi‐omics integration, enabling robust identification of structured molecular dependencies across diseases.

Minghui Zhao, Na Zhou, Ruotong Liu et al. · 0 citations
Open access Aug 2026

iDCF: Interpretable deconvolution of cell fractions via biologically-informed deep learning using scRNA-seq data.

iDCF (Interpretable Deconvolution of Cell Fractions) is a novel framework that enforces biological topology onto deep neural networks, bridging the gap between computational inference and biological intuition.

Hongming Guo, Tingfang Wu, Wen-Zheng Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.