Abstract P15: Context-aware foundation model of bulk transcriptomics for interpretable analysis of transcriptional dynamics and treatment response in AML
A context-aware foundation model of bulk transcriptomics for interpretable analysis of transcriptional dynamics and treatment response in AML, which establishes the model as a powerful framework to extract useful information from bulk transcriptomics data and has potential applications in precision medicine by connecting computational predictions with biological insights.
Abstract
While language models extract linguistic structures from text, similar approaches can uncover biological rules from genetic patterns. Though these methods have shown promise in single-cell analysis, bulk transcriptomics remains underexplored despite offering distinct clinical advantages including preserved tissue-level information, higher sequencing depth, and cost-effectiveness. Here, we present a transformer-based foundation model leveraging transcriptomic profiles from over 30,000 diverse bulk RNA samples, including normal tissues and various cancer types. Unlike conventional language models, our model incorporates specialized modules for modelling pairwise gene interactions through a dual representation system that captures both gene-level features and their higher-order relationships. Our model shows robust performance across multiple downstream applications. It achieves zero-shot accuracy of 78.81% in cancer classification without fine-tuning and outperforms existing approaches in cancer stages prediction through simple fine-tuning. Notably, it can extract critical gene interaction networks without relying on prior biological knowledge. More importantly, we leverage it to introduce dynamic interpretations to static bulk transcriptomic data, successfully modelling logical gene regulation rules with 91.07% overall accuracy—reaching 100% for rules related to key genes like GATA2 and SCL. With the context-specific modelling ability, it also identifies, for example, transcriptional dynamics in normal haematopoiesis and dysregulated circuits during transition to leukemic states. We further demonstrate clinical utility in predicting patient response to first induction chemotherapy (AUROC=0.75) in acute myeloid leukemia, a challenging task due to patient and mechanism heterogeneity. Through our novel response-directed feature-space gradient ascent approach, we identify patient-specific gene expression modifications that could computationally redirect resistant phenotypes toward responsive ones, revealing potential therapeutic targets aligned with individual patients' clinical features. These results establish our model as a powerful framework to extract useful information from bulk transcriptomics data and has potential applications in precision medicine by connecting computational predictions with biological insights.
Yi Chai, Yang Li, Jianbiao Zhou, Wee Joo Chng, Yang Zhang. Context-aware foundation model of bulk transcriptomics for interpretable analysis of transcriptional dynamics and treatment response in AML [abstract]. In: Proceedings of Frontiers in Cancer Science 2025; 2025 Nov 5-7; Singapore. Philadelphia (PA): AACR; Cancer Res 2026;86(13_Suppl):Abstract nr P15.
Single-cell RNA sequencing (scRNA-seq) enables detailed characterization of cellular heterogeneity, yet understanding the full cellular and regulatory environment of complex tissues remains challenging. In the era of large single-cell atlases, this technology has become increasingly accessible, and datasets have grown in scale and statistical power. As a result, sample representation methods have emerged as a promising strategy to summarize patient-level biological variation. However, most existing approaches rely on unsupervised learning frameworks with ambiguous biological interpretability. Here we present a Framework for Learning Over REgulatory-Embedding Networks (FloREN), a supervised and interpretable sample representation method. FloREN models single-cell data as a heterogeneous network integrating cells and genes together with gene regulatory and cell-cell communication relationships. Through condition-aware embeddings and interpretable attention networks, FloREN enables improved sample stratification and biomarker discovery. In addition, the framework supports downstream analyses that found specific immune network mechanisms in immune-mediated inflammatory diseases (IMIDs).
Iñigo Clemente‐Larramendi, S. Hillion, D. Cornec et al.· bioRxiv· 0 citations
Traditional RNA-seq analysis depends heavily on genome alignment and gene annotation, limiting its utility in non-model organisms and introducing biases that can obscure regulatory complexity. We present MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer–based statistics to detect biological variation without requiring alignment. MKMC integrates fast k-mer counting, abundance matrix generation, normalization, dimensionality reduction, and differential analysis into a unified workflow. Across diverse datasets, MKMC recapitulates key biological signals—including sex differences in killifish liver—and matches alignment-based pipelines in differential expression analysis and transcriptomic age prediction. Notably, MKMC detects isoform-specific events missed by traditional methods, one of which we validated using in situ hybridization. These results reveal previously hidden isoform-level regulatory events that contribute to sex-and age-associated transcriptional programs. MKMC offers a robust, extensible alternative to alignment-based approaches, enabling transcriptomic discovery across both model and non-model systems. While we focus here on RNA-seq as a primary application, MKMC is broadly applicable to any k-mer–based analysis of next-generation sequencing data.
L. Mboning, Maciej Dlugosz, Marek Kokot et al.· bioRxiv· 0 citations
Current liquid biopsy methods for multi-cancer detection using plasma cell-free RNA (cfRNA, short RNA fragments circulating in blood that can reflect disease states) typically rely on gene annotations, which can overlook signals from unannotated or repetitive genomic regions. We present GeneLLM, a Transformer-based model that directly processes the nucleotide sequences of human-mapped cfRNA reads to identify cancer-indicative signatures. By bypassing gene-level quantification, the model retains signals from transcriptomic dark matter. The model learns latent pseudo-biomarkers (prototype representations from aggregated cfRNA read embeddings) that serve as discriminative features for cancer classification, rather than corresponding to explicit genomic sequences. Here we show that, in a multi-centre cohort, GeneLLM achieves ROC-AUC values ranging from 0.9250 to 0.9962 across several cancers, while maintaining comparable performance at one-sixth of the typical sequencing depth. These results suggest that sequence-level modelling of plasma cfRNA can capture diagnostically relevant information beyond annotation-dependent approaches, enabling more cost-efficient and scalable cancer screening. Cell-freeRNA (cfRNA) can be a non-invasive and cost-effective biomarker for cancer therapy and clinical outcomes, but its analysis remains challenging. Here, the authors develop GeneLLM, a cfRNA-based large language model that processes raw cfRNA data and allows accurate cancer classification from plasma biopsies.
Siwei Deng, Lei Sha, Yongcheng Jin et al.· Nature Communications· 0 citations
Genomic foundation models are increasingly reused as frozen feature extractors for downstream sequence prediction, offering a compute-efficient alternative to full fine-tuning. However, it remains unclear when biological information encoded by these models is accessible without task-specific adaptation. We present a representation-accessibility analysis of frozen genomic language models across regulatory, epigenetic, promoter, splice-site, and variant-effect prediction tasks. We evaluate DNABERT-2, Nucleotide Transformer, HyenaDNA, GENERATOR-v2, and Omni-DNA under unified frozen-probing protocols, while separating diagnostic readout analyses from validation-selected checks. Our results reveal a consistent task-dependent pattern: frozen probes recover 95-100 % of fine-tuned performance on promoter tasks, but average splice-site recovery drops to 60-88 %. Frozen embeddings are also competitive on broad Genomic Benchmark tasks such as coding-region and species-discrimination classification, but show larger gaps on some regulatory and OCR tasks. Layer-wise probing, in-silico mutagenesis, variant-effect prediction, and embedding geometry show that local biological signal is partially present in frozen representations, but is not always accessible through final pooled embeddings.
Nirjhor Datta, Swakkhar Shatabda, M. S. Rahman· 0 citations
Abstract Motivation Gene set analysis (GSA) is a foundational approach for interpreting genomic data of diseases by linking genes to biological processes. However, conventional GSA methods overlook clinical context of the analyses, often generating long lists of enriched pathways with redundant, nonspecific, or irrelevant results. Interpreting these requires extensive, ad-hoc manual effort, reducing both reliability and reproducibility. Results We introduce cGSA, a novel AI-driven framework that enhances GSA by incorporating context-aware pathway prioritization. cGSA integrates gene cluster detection, enrichment analysis, and large language models to identify pathways that are not only statistically significant but also biologically meaningful. Benchmarking on 102 curated gene sets across 19 diseases and ten disease-related biological mechanisms shows that cGSA outperforms baseline methods by over 30%, with expert validation confirming its increased precision and interpretability. Two independent case studies in melanoma and breast cancer further demonstrate its potential to uncover context-specific insights and support targeted hypothesis. Availability and Implementation The demo website is publicly available at https://www.ncbi.nlm.nih.gov/CBBresearch/Lu/Demo/cGSA/, while the data and code can be accessed at https://github.com/ncbi-nlp/cGSA.
Abstract Single-cell RNA sequencing has emerged as a transformative tool, enabling precise phenotype prediction and the detailed identification of disease-associated cell subpopulations. However, many existing computational approaches still rely on predefined cell-type annotations during model training. This dependence makes their predictive performance highly sensitive to subjective annotation quality, labeling inconsistencies, and dataset-specific biases, ultimately hindering their generalizability across diverse patient cohorts. To address these challenges, we propose scCap, an annotation-free framework that leverages knowledge-augmented clustering for robust phenotype prediction. Specifically, the framework first constructs initial clusters from raw gene expression profiles and subsequently refines them within the embedding space of a pretrained single-cell foundation model, allowing the clusters to better reflect broader biological organization while preserving fine-grained cellular heterogeneity. The resulting knowledge-augmented clusters are then integrated into a hierarchical multiple instance learning framework with dual-level attention, enabling interpretable predictions at both the cell and cluster levels. Evaluated across three public scRNA-seq datasets, scCap consistently outperforms baseline models in predictive accuracy. Furthermore, scCap identifies disease-associated subpopulations previously reported in the literature without relying on predefined cell-type annotations. These results demonstrate that scCap provides a robust and interpretable framework for annotation-free phenotype prediction.
Janghyun Noh, Yoobin Shin, M. Kim et al.· Briefings in Bioinformatics· 0 citations