Evaluation of Self-Supervised Pre-training for Disease Classification from Gut Metagenomes
Abstract
Self-supervised pre-training has become the dominant paradigm in single-cell and protein modelling, yet its value for microbiome-based disease classification remains untested under evaluation protocols that control for study-level confounding. Public metagenomic collections aggregate hundreds of independently conducted studies, and disease labels are strongly confounded with study identity, so a classifier evaluated on random splits may recognize the contributing cohort rather than the disease. We constructed a benchmark from curatedMetagenomicData comprising three independent binary healthy-versus-disease tasks: colorectal cancer (CRC; 1,395 samples, 701 cases, 11 studies), inflammatory bowel disease (IBD; 448 samples, 185 cases, 2 studies) and type 2 diabetes (T2D; 1,632 samples, 773 cases, 3 studies). Three controls were imposed: every subject contributing to a fine-tuning cohort was excluded from the pre-training corpus, all data splits were grouped by subject, and eight cohort- and sequencing-level covariates acting as study proxies were removed from every feature set. A feature-tokenizer transformer encoder was pre-trained by masked feature reconstruction and fine-tuned across a factorial grid of three feature arms, two encoder depths, three embedding dimensions and five random seeds, yielding 216 result files and 3,420 model fits. Each pre-trained configuration was matched to a from-scratch counterpart, and three classical baselines were fitted on identical folds and seeds. Taking the held-out cohort as the unit of inference, pre-training produced a positive but statistically non-significant improvement in cross-cohort discrimination: Δ macro-AUC = +0.011 across 16 cohorts (95% CI −0.005 to +0.027; 10 of 16 cohorts improved; Wilcoxon signed-rank p = 0.21). The effect was concentrated in the diseases with least cohort replication, IBD (+0.032, 2 cohorts) and T2D (+0.030, 3 cohorts), and was absent in CRC (+0.002 across 11 cohorts, 6 of 11 improved). A secondary pattern was significant: under leave-one-study-out evaluation, pre-trained encoders were less sensitive to embedding dimension than from-scratch encoders (initialization-by-capacity interaction +0.033 macro-AUC, 95% CI +0.003 to +0.062, p = 0.021), with no corresponding effect under within-cohort evaluation. Classical baselines using default hyperparameters matched or exceeded the transformer in all cells of a pre-specified comparison. Cross-study heterogeneity exceeded every model-level difference observed: per-study CRC macro-AUC ranged from 0.44 to 0.75 across the 11 held-out cohorts. We conclude that at currently available public corpus sizes, self-supervised pre-training does not confer a reliable cross-cohort advantage for metagenomic disease classification, and that cohort heterogeneity rather than model architecture constitutes the binding constraint. The benchmark, pre-trained weights, cohort definitions and all result files are released to permit re-evaluation against larger corpora.