Skip to content
#protein folding Open access

Evaluation of Self-Supervised Pre-training for Disease Classification from Gut Metagenomes

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)
Single-cell and spatial transcriptomics

Abstract

Self-supervised pre-training has become the dominant paradigm in single-cell and protein modelling, yet its value for microbiome-based disease classification remains untested under evaluation protocols that control for study-level confounding. Public metagenomic collections aggregate hundreds of independently conducted studies, and disease labels are strongly confounded with study identity, so a classifier evaluated on random splits may recognize the contributing cohort rather than the disease. We constructed a benchmark from curatedMetagenomicData comprising three independent binary healthy-versus-disease tasks: colorectal cancer (CRC; 1,395 samples, 701 cases, 11 studies), inflammatory bowel disease (IBD; 448 samples, 185 cases, 2 studies) and type 2 diabetes (T2D; 1,632 samples, 773 cases, 3 studies). Three controls were imposed: every subject contributing to a fine-tuning cohort was excluded from the pre-training corpus, all data splits were grouped by subject, and eight cohort- and sequencing-level covariates acting as study proxies were removed from every feature set. A feature-tokenizer transformer encoder was pre-trained by masked feature reconstruction and fine-tuned across a factorial grid of three feature arms, two encoder depths, three embedding dimensions and five random seeds, yielding 216 result files and 3,420 model fits. Each pre-trained configuration was matched to a from-scratch counterpart, and three classical baselines were fitted on identical folds and seeds. Taking the held-out cohort as the unit of inference, pre-training produced a positive but statistically non-significant improvement in cross-cohort discrimination: Δ macro-AUC = +0.011 across 16 cohorts (95% CI −0.005 to +0.027; 10 of 16 cohorts improved; Wilcoxon signed-rank p = 0.21). The effect was concentrated in the diseases with least cohort replication, IBD (+0.032, 2 cohorts) and T2D (+0.030, 3 cohorts), and was absent in CRC (+0.002 across 11 cohorts, 6 of 11 improved). A secondary pattern was significant: under leave-one-study-out evaluation, pre-trained encoders were less sensitive to embedding dimension than from-scratch encoders (initialization-by-capacity interaction +0.033 macro-AUC, 95% CI +0.003 to +0.062, p = 0.021), with no corresponding effect under within-cohort evaluation. Classical baselines using default hyperparameters matched or exceeded the transformer in all cells of a pre-specified comparison. Cross-study heterogeneity exceeded every model-level difference observed: per-study CRC macro-AUC ranged from 0.44 to 0.75 across the 11 held-out cohorts. We conclude that at currently available public corpus sizes, self-supervised pre-training does not confer a reliable cross-cohort advantage for metagenomic disease classification, and that cohort heterogeneity rather than model architecture constitutes the binding constraint. The benchmark, pre-trained weights, cohort definitions and all result files are released to permit re-evaluation against larger corpora.

View source

Similar papers

#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Book Open access Jul 2015

Understanding the affect of developers: theoretical background and guidelines for psychoempirical software engineering

This paper highlights the challenges to conduct proper affect-related studies with psychology, provides a comprehensive literature review in affect theory, and proposes guidelines for conducting psychoempirical software engineering.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 56 citations · ⚡4
#machine learning Open access May 2017

What Influences the Speed of Prototyping? An Empirical Investigation of Twenty Software Startups

This study conducts a multiple case study on twenty European software startups and proposes a prototype-centric learning model in early stage software startups, and identifies factors that occur as barriers but also facilitators for prototyping in earlystage software startups.

Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson · 44 citations · ⚡5
#protein folding Open access Sep 2026

Programmable design of functional proteins from natural language

Pinal, a 16-billion-parameter foundation model that produces protein candidates from natural-language functional descriptions, supports natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or seque...

Fengyuan Dai, Shiyang You, Yudian Zhu et al. · 31 citations · ⚡3

Related blog posts

Google DeepMind Blog Sep 30, 2026

Introducing SynthID Bio

Proof of concept for watermarking AI-generated proteins while preserving biological function.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.