Phenoverse is introduced, an interpretable deep learning framework that learns sample-level disease state representations through cell type-aware residual encoding, prototype learning, and Perceiver-based aggregation that demonstrates that trajectory-derived genes reveal cross-cohort molecular programs and show consistently higher reproducibility than traditional case-control comparisons.
Abstract
Single-cell transcriptomics technology offers unprecedented insights into molecular heterogeneity. However, capturing sample-level representations that reflect both systemic and cellular states remains challenging, especially when disease annotations are mostly available as coarse sample-level labels. Here, we introduce Phenoverse, an interpretable deep learning framework that learns sample-level disease state representations through cell type-aware residual encoding, prototype learning, and Perceiver-based aggregation. Applied to independent single-cell transcriptomic cohorts of COVID-19, Alzheimer’s disease, and systemic lupus erythematosus, totaling over 5 million cells, we demonstrate that learned sample representations enable disease state prediction and encode a continuous spectrum of disease severity on unseen data that correlate with multiple clinical and pathological measures, despite being trained solely on binary phenotype labels. Further, we demonstrate that trajectory-derived genes reveal cross-cohort molecular programs and show consistently higher reproducibility than traditional case-control comparisons. Finally, prototype learning provides intrinsic model interpretability and enables the characterization of cell type-specific disease states. Taken together, Phenoverse offers an interpretable disease-phenotyping approach to dissecting sample heterogeneity, and our results highlight its utility in translating complex single-cell transcriptomic data into patient-level biological insights.
A Framework for Learning Over REgulatory-Embedding Networks (FloREN), a supervised and interpretable sample representation method that enables improved sample stratification and biomarker discovery and supports downstream analyses that found specific immune network mechanisms in immune-mediated inflammatory diseases (IMIDs).
Iñigo Clemente‐Larramendi, S. Hillion, D. Cornec et al.· bioRxiv· 0 citations
GITIII-scale is presented, a hierarchical, interpretable pan-cancer spatial transcriptomics foundation model for TME representation learning that investigates cell state-niche associations and their underlying ligand-receptor (LR) signaling pathways that recovered niche-associated state changes more accurately than existing spatial transcriptomics foundation models in cancer types unseen during training.
Xiaohui Xiao, Jia-Shu He, Shiyang Zhang et al.· 0 citations
Spatial transcriptomics now profiles patient cohorts at single-cell resolution, enabling analysis of disease-associated cell organization in situ. However, discovering such spatial biomarkers remains challenging because relevant structures occur at unknown scales and cell-or niche-level annotations are rarely available. We present spHOT, a deep learning framework that localizes phenotype-associated spatial biomarkers from sample-level labels. spHOT combines spatial foundation model embeddings, a hierarchical domain tree for multi-resolution tissue representation, and a teacher-student multiple instance learning architecture that converts sample labels into cell-level biomarker scores. In controlled simulations and real-tissue benchmarks, spHOT outperformed existing spatial and single-cell methods in localizing ground-truth biomarkers. Across fibrotic, metabolic, and autoimmune disease datasets, spHOT recovered disease-relevant niches and tissue states reported by supervised analyses in the original studies. Cross-disease application of spHOT transferred biomarkers across chronic lung diseases without retraining. spHOT enables scalable, annotation-efficient spatial biomarker discovery in cohort-scale spatial transcriptomics.
H. Kim, Donghee Kim, Sangwook Jung et al.· bioRxiv· 0 citations
Identifying small, interpretable gene sets that robustly capture disease-associated variation in singlecell transcriptomic data remains a central challenge for biological interpretation and experimental followup. In practice, commonly used differential expression and sparsity-based approaches often produce large, unstable gene lists that fail to generalize across patients due to strong donor-specific confounding. We study sparse gene selection for reconstructing donor-robust, disease-aligned cellular trajectories in real single-cell RNA-seq datasets. We introduce Sparse Linear Manifold Control (SLMC), a practical workflow that defines a disease-aligned score after removing donor-associated variation and selects minimal gene programs whose expression reconstructs this score. We focus on diagnosing the structure of the resulting reconstruction objective and evaluating selection strategies under realistic health data conditions. Across five human single-cell datasets spanning oncology and neurodegeneration, we find that the reconstruction objective exhibits strong diminishing returns, explaining why simple greedy selection methods perform well in practice. Under strict donor-heldout evaluation, greedy methods consistently outperform LASSO at small gene budgets and achieve accurate reconstruction with as few as 25 genes. Together, these results highlight how careful objective design and empirical evaluation enable robust and interpretable gene selection for disease-aligned representation learning in single-cell health data. Data and Code Availability All code used for data preprocessing, model training, evaluation, and feature importance analyses is available at: https://github.com/AdiVM/SLMC_single-cell. The singlecell data used in this study are publicly available human transcriptomic datasets generated by prior studies and accessible through the Gene Expression Omnibus (GEO). Analyses were performed using renal cell carcinoma single-cell RNA-seq data (GEO accession: GSE314072) and Alzheimer’s disease single-nucleus RNA-seq data from human cortex (GEO accession: GSE138852), with additional publicly available datasets used for cross-context diagnostic evaluation (GEO accessions: GSM8652069, GSE308624, and GSE227734). All datasets contain de-identified human samples and are available under standard publicuse terms via GEO. Institutional Review Board (IRB) This study analyzes de-identified, publicly available human transcriptomic data obtained from previously published studies. No new data were collected, and no identifiable private information was accessed. In accordance with institutional policy, this work was determined to constitute non-human subjects research and did not require additional IRB approval.
Adithya V. Madduri, Chirag J. Patel· bioRxiv· 0 citations
In biomarker discovery, access to sufficient quantities of condition-specific transcriptomic data is often limited by cohort size, privacy concerns, and domain shift between normal and condition populations. Generative modeling can augment scarce cohorts and probe distributional transitions. Furthermore, synthetic transcriptome generation can support differential expression analyses, machine learning, privacy-preserving data sharing, benchmarking, and hypothesis generation in translational bioinformatics workloads in fields such as oncology. Here, we present Hoike, a framework that combines a crossdomain Joint-Embedding Predictive Architecture (JEPA) with a latent diffusion model to generate condition-specific bulk transcriptomes from a normal reference context. In Hoike, normal tissue profiles provide continuous conditioning signals, while the model learns disease-linked shifts in latent space and reconstructs gene-level expression in log2(TPM+1) space. The implementation supports paired normal-condition training, tissuealigned conditioning, and constrained non-negative decoding for biologically valid outputs. We describe the architecture, objective design, and evaluation protocol used in this work across GTEx-derived normal references and multiple TCGA condition cohorts as a case study. This serves as the technical specification of the Hoike framework and its reproducible analysis workflow.
Artificial Intelligence (AI) has been increasingly applied to investigate genetic irregularities associated with Alzheimer's Disease (AD). However, its potential to uncover deeper, more detailed molecular and cellular mechanisms remains underexplored, primarily due to limitations in integrating large-scale data and capturing the complex, cell-type-specific dynamics involved in AD pathology. Single-cell RNA sequencing (scRNA-seq) has emerged as a powerful tool in transcriptomics, offering high-resolution, cell-specific insights into complex biological systems. Despite this advancement, a significant gap remains in identifying both common and cell-type-specific transcriptomic signatures that define AD-related cellular and molecular processes. To address this, we propose a deep learning framework leveraging a Multi-Layer Perceptron (MLP) to classify AD versus control nuclei using scRNA-seq data from the Religious Orders Study/Memory and Aging Project (ROSMAP). We focus on microglial subclusters, particularly those representing homeostatic and activated states, to train the MLP model for optimal classification performance. We utilize the predicted embeddings from the MLP to model a disease progression trajectory for each of the datasets. Our model demonstrates strong performance in both classification and disease trajectory inference. To enhance interpretability, SHapley Additive exPlanations (SHAP) are applied to identify key AD-associated genes. Based on the most salient genes implicated in AD, we built transcription gene regulatory networks, revealing novel transcription factors and regulons for AD pathogenesis. These regulons highlight profound impacts of dysregulations of proteostasis, endoplasmic reticulum (ER) stress responses, and circadian rhythm on synaptic plasticity and neuronal survival in AD, offering a more holistic approach to drug target discovery compared to conventional single-target strategies, potentially leading to greater efficacy in slowing or reversing disease progression. This work demonstrates the transformative potential of AI in elucidating the molecular mechanisms of AD, offering improvements over traditional methods and uncovering novel insights into disease pathogenesis and potential therapeutic targets.
M. Trivedi, Jay Shah, Yi Su et al.· AI in neuroscience· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.