VirTues-derived biomarkers predict anti-PD-L1 chemo-immunotherapy response and stratify disease-free survival in an independent cohort, outperforming state-of-the-art biomarkers derived from the same datasets and current clinical stratification schemes.
Abstract
Spatial proteomics technologies have transformed our understanding of complex tissue architecture in cancer but present unique challenges for computational analysis1. Each study uses a different marker panel and protocol, and most methods are tailored to single cohorts, which limits knowledge transfer and robust biomarker discovery. Here we present Virtual Tissues (VirTues), a general-purpose foundation model for spatial proteomics that learns marker-aware, multi-scale representations of proteins, cells, niches and tissues directly from multiplex imaging data. From a single pretrained backbone, VirTues supports marker reconstruction, cell segmentation and typing, niche annotation, spatial biomarker discovery and patient stratification, including zero-shot annotation across heterogeneous panels and datasets. In triple-negative breast cancer, VirTues-derived biomarkers predict anti-PD-L1 chemo-immunotherapy response2 and stratify disease-free survival in an independent cohort3, outperforming state-of-the-art biomarkers derived from the same datasets and current clinical stratification schemes.
Spatial transcriptomics now profiles patient cohorts at single-cell resolution, enabling analysis of disease-associated cell organization in situ. However, discovering such spatial biomarkers remains challenging because relevant structures occur at unknown scales and cell-or niche-level annotations are rarely available. We present spHOT, a deep learning framework that localizes phenotype-associated spatial biomarkers from sample-level labels. spHOT combines spatial foundation model embeddings, a hierarchical domain tree for multi-resolution tissue representation, and a teacher-student multiple instance learning architecture that converts sample labels into cell-level biomarker scores. In controlled simulations and real-tissue benchmarks, spHOT outperformed existing spatial and single-cell methods in localizing ground-truth biomarkers. Across fibrotic, metabolic, and autoimmune disease datasets, spHOT recovered disease-relevant niches and tissue states reported by supervised analyses in the original studies. Cross-disease application of spHOT transferred biomarkers across chronic lung diseases without retraining. spHOT enables scalable, annotation-efficient spatial biomarker discovery in cohort-scale spatial transcriptomics.
H. Kim, Donghee Kim, Sangwook Jung et al.· bioRxiv· 0 citations
X-SPATIO is a spatially compatible computational pipeline designed to directly link hematoxylin and eosin morphology with region-matched mRNA and protein expression, enabling cost-effective inference of spatial biomarker expression and establishing a foundation for biologically grounded discovery and precision oncology in TNBC.
Vibha R. Rao, Madhumala K. Sadanandappa, C. Black et al.· American Journal of Patholog...· 0 citations
Spatial proteomics (SP) measures the spatial distribution of proteins within tissues, providing important insights into tissue function, disease, and therapeutic response. However, current SP technologies profile only a small fraction of the proteome and are limited by cost and measurement noise. Recent AI approaches enable predicting spatial protein expression from transcriptomic or histopathological data, but are typically restricted to paired datasets covering only tens of proteins, limiting their ability to generalize beyond experimentally measured protein panels. Here we present SPgen, a multi-modal foundation-model framework for proteome-wide spatial protein prediction. SPgen integrates protein sequences, functional annotations, transcriptomic profiles, and spatial information to learn transferable representations that enable inference beyond experimentally profiled proteins. Across diverse spatial proteomics datasets, SPgen accurately reconstructs measured spatial patterns, reduces measurement noise, and enables proteome-wide spatial prediction.
Jiachen Li, Kaiyuan Yang, Qiaoling Che et al.· bioRxiv· 0 citations
Spatial transcriptomics reveals cellular heterogeneity, intercellular communication, and tissue organization, but its cost and limited accessibility restrict clinical use. Here, we present VISTA, a model that integrates multi-scale histological features and spatial context to infer spatial gene expression from H&E-stained tissue images. Across leave-one-section-out cross-validation and independent validation, VISTA robustly predicted thousands of genes and outperformed state-of-the-art methods. Beyond expression reconstruction, VISTA enabled clinically relevant downstream analyses. In TCGA breast cancer samples, it identified survival-associated genes, stratified prognostic risk groups, and revealed adverse tumor-associated spatial subtypes. In our in-house intrahepatic cholangiocarcinoma cohort, it preserved tumor–normal organization and identified CLDN4 and CYP3A4 as complementary spatial biomarkers. In HER2+ breast cancer, it predicted pathological response to neoadjuvant trastuzumab-based therapy and linked response-associated regions to immune and cytokine-related programs. These results support virtual spatial transcriptomics from routine histopathology for oncology applications.
Shaoqing Jiao, Zhen Yuan, Dazhi Lu et al.· bioRxiv· 0 citations
Spatial omics (SO) technologies enable spatially resolved molecular profiling, while hematoxylin and eosin (H&E) imaging remains the gold standard for morphological assessment in clinical pathology. Recent computational advances increasingly place H&E images at the center of SO analysis, bridging morphology with transcriptomic, proteomic, and other spatial molecular modalities. This lecture-style tutorial surveys the algorithmic foundations and recent advances in method development that make them practical and impactful for precision medicine. Following the tutorial flow, we first introduce key SO modalities and data abstractions (tiles/patches, spots, cells, and spatial graphs) and articulate problems to address and motivations, emphasizing multi-scale mismatch, structured spatial dependence, weak supervision, and domain shift across cohorts and sites. We then trace the evolution of modern multimodal representation learning, highlighting graph neural networks, transformer-based architectures, and encoder–decoder designs. The core of the tutorial systematically organizes contemporary methods into three categories: (i) integration methods, which jointly model paired multimodal measurements; (ii) mapping methods, which predict spatial molecular profiles from H&E images; and (iii) foundation models (FMs), which learn transferable representations from large-scale spatial datasets via self-supervised pretraining, contrastive objectives, etc. This tutorial also discusses applications of generative modeling to support imputation and data augmentation. Throughout, we connect methods to real biomedical endpoints (e.g., tumor microenvironment characterization, biomarker discovery, and cohort-level stratification). We further summarize actionable modeling directions enabled by current architectures and delineate persistent gaps driven by data, biology, and technology that are unlikely to be resolved by model design alone. The tutorial concludes with open challenges in interpretability, reliability, privacy, and clinical translation, outlining opportunities for the KDD community to contribute principled data mining and learning approaches to multimodal spatial biology.
Ninghui Hao, Boshen Yan, Dong Li et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.