Skip to content
Preprint

CytoBERT: A Foundation Model for Cytometry Data

Aug 2026 · 0 citations · 30 references
Computer Science

TL;DR

Fine-tuning CytoBERT for sample-level classification demonstrates that transfer learning across heterogeneous cytometry datasets is feasible, providing a starting point for scalable, generalizable cytometry analysis.

Abstract

Cytometry measures the complex characteristics of single cells (e.g., counts and protein expression of immune cells) and is widely used across immunological research and clinical settings. However, cytometry data is highly heterogeneous and unstandardized due to experimental protocols and the choice of measured features. While machine learning methods hold the potential to gain deeper insights into cell biology, these challenges make them difficult to apply and transfer across studies. Recent advances in foundation models can alleviate these issues, but corresponding approaches are still scarce in this field. To address this, we provide CytoBERT, a publicly available, open-source, open-weight foundation model for single-cell cytometry data with variable marker panels. CytoBERT is pretrained in a self-supervised manner on a large-scale cytometry corpus (15 human datasets with heterogeneous marker panels and more than 50 million cells) curated through marker standardization, enabling it to learn transferable inter-marker relationships within cells. Fine-tuning CytoBERT for sample-level classification demonstrates that transfer learning across heterogeneous cytometry datasets is feasible, providing a starting point for scalable, generalizable cytometry analysis. Code is available at GitHub.

View source

Similar papers

Open access Sep 2026

Velociraptor Machine Learning Quantifies Similarity to Known Cell Types and Matches Cells Across Flow and Imaging Cytometry Platforms.

Suspension flow cytometry enables high-throughput cellular profiling at the single cell level, but these data lack positional information. Conversely, tissue-based imaging cytometry techniques reveal a cell's location within the tissue architecture and can provide insight into cell biology. It would be especially valuable if data analysis tools could incorporate data from imaging and flow cytometry platforms to gain complementary strengths when quantifying features of cells and populations. We hypothesized that per-cell Marker Enrichment Modeling (MEM) might provide a way to register cells between flow and imaging cytometry analysis. Here, we developed the Velociraptor machine learning workflow for cross-platform cytometry analysis. Velociraptor begins with a graph-based implementation of MEM to calculate per-cell quantitative phenotype labels. With this information, Velociraptor can then quickly calculate similarity between each cell's phenotype and search terms describing established cell types, cells of interest, or cells observed in other samples. Velociraptor was effective in registering cells within and between cytometry platforms. Integrated identification of cell populations was tested in several challenges, including comparisons of high dimensional datasets from cancer and immunology. Tested instrument types included imaging mass cytometry (IMC), cyclic immunohistochemistry (cycIHC), suspension mass cytometry (CyTOF), and suspension spectral flow cytometry (SFC). Between IMC and CyTOF, a comparison across imaging and flow cytometry platforms that use the same mass tag probes, Velociraptor accurately identified and registered immune cell types (median F1-measure of 0.81). Between SFC and CyTOF, a comparison of two fundamentally different probe types-fluorophores and metal tags-in suspension flow cytometry, Velociraptor was even more accurate at identifying and registering cells (concordance correlation coefficient of 0.99). Velociraptor was especially useful in heterogeneous samples where individual cells diverged in phenotype from the bulk population. In IMC imaging of human breast cancer, a previously unappreciated tumor cell subset was revealed by Velociraptor, characterized as CD15+, and validated as spatially segregated to the tumor core. In both cycIHC (8-dimensional imaging) and IMC imaging (40-dimensional imaging), Velociraptor accurately identified macrophages using a single search label as input. Notably, Velociraptor worked effectively with both extremely rare and highly abundant cell types and with cell search labels calculated from data and theoretical labels based on literature and expertise. The Velociraptor algorithm is freely available at https://github.com/cytolab.

Claire E. Cross, Asa A. Brockman, Rebecca A. Ihrie et al. · 0 citations
#small language model Open access Aug 2026

CytoGate-Bench: an LLM benchmark for cross-panel cell gating in cytometry

This work introduces CytoGate-Bench, a benchmark that reformulates this per-step procedure as a zero-shot, panel-agnostic task for large language models, and contributes a public benchmark that tests precisely that ability across 11 human cohorts.

Jaesik Kim, Byounghan Lee, Namhyuk Ahn et al. · 0 citations
Open access Jul 2026

Optimal transport analysis of high-dimensional flow cytometry data in immuno-oncology

Introduction Advances in single-cell and spatial profiling have enabled detailed characterization of heterogeneous samples, but analyzing this data remains challenging in settings involving multiple comparisons. While tools like UMAP and t-SNE are valuable for visualization, their stochastic, parameter-sensitive nature limits their use in longitudinal comparisons, treatment group analysis, and multicenter trials. Although OT was first described in the 19th century, the Sinkhorn algorithm makes it computationally tractable for high-dimensional data. By directly comparing distributions of cellular states, OT provides reproducible measures of change in high-dimensional space. This framework is amenable to integration with machine learning, including deep generative models. Methods OT was applied to longitudinal data from a phase I trial of tocilizumab for cavitary malignancies (NCT 06016179). The current implementation makes use of expert-guided phenotypic population definitions and their relationships. An OT-based graph representation was created for baseline and follow-up samples. The graph layout was fixed across samples and computed from phenotypic relationships. In this implementation vertex radii are proportional to their relative abundance, allowing for rapid visual assessment of population-level increases and decreases. Graph edge thickness and color encode inter-population similarity based on the optimal transport (Sinkhorn) distance between marker expression distributions. Results This representation enabled rapid identification of populations undergoing substantial change, such as the CD8+/IFNɣ+ population, which decreased from 63% to 17% of CD8+ T cells following treatment. Population changes across all fluorescence parameters were encoded in the graph edit distance (GED), which captures changes in population abundance and phenotypic shifts in marker space. Discussion Future implementations can combine this expert-guided approach with unbiased clustering algorithms to enhance scalability and cross-platform harmonization. In our recently initiated clinical trials, we will apply OT to identify key shifts in tumor, immune, and stromal cell states, summarizing patient trajectories and quantitatively supporting predictive models of treatment response. Potential applications include quantifying residual disease after chemotherapy, tracking immune activation during immunotherapy, and linking host–microbiome interactions to disease progression. This approach overcomes the limitations of traditional, local-structure-optimized tools (UMAP or t-SNE) to provide a comprehensive, longitudinal view of tumor evolution and treatment response.

Abida Sanjana Shemonti, Justin C. Wang, Albert D. Donnenberg et al. · 0 citations
Jul 2026

A 61-Parameter CyTOF Panel for Comprehensive Profiling of Human PBMC to Characterize Activation, Differentiation, Checkpoints and Cytokines 2259198

High-parameter cytometric analysis enables discoveries of potential therapeutic targets and informs disease prognoses in translational and clinical research. CyTOF™ technology is a single-cell analysis platform that uses metal-tagged antibodies to resolve 50-plus markers in a single tube. Notably, antibody cocktails and stained samples can be frozen for later use and acquisition, or barcoded and pooled together, minimizing technical variation. Unlike fluorescence-based cytometry, spectral unmixing and single-stain controls are not required. Therefore, it is uniquely possible with CyTOF technology to rapidly design high-parameter panels and use intracellular markers to gain functional insights. The goal of this study was to design a 61-parameter CyTOF panel for in-depth functional immune profiling of human PBMC. The panel contains lineage markers for major immune cell subsets and a diverse array of phenotyping T cell targets focused on activation, differentiation, checkpoint and cytokines and can be used to identify over 60 cell populations. Untreated and stimulated PBMC were barcoded, pooled and stained. Samples were frozen and acquired on a later day using a CyTOF XT PRO system. High-dimensional analysis of stimulated PBMC revealed striking cross-lineage immuno-functional diversity at the single-cell level. The expression of over 20 markers spanning the functional landscape of activation, checkpoints and cytokines was revealed across effector, memory, cytotoxic, regulatory and exhausted immune cell populations. CyTOF systems enable the highest number of simultaneous measurements in a single panel, allowing for wide immune coverage and high resolution of intracellular targets to interrogate functional potential. Overall, studying functional immunology using CyTOF technology can elucidate the complex nature of immune responses to provide an understanding of how they relate to disease and treatments. For Research Use Only. Not for use in diagnostic procedures., n/a Immune Mechanisms of Human Disease (HUM)

Michael J. Cohen, Stephen Li, Lauren J. Tracey et al. · 0 citations
Review Open access Aug 2026

CytoFormer: A Molecularly Supervised Cell Foundation Model for Histopathology Cell Classification

Identifying cell types directly from routine haematoxylin and eosin (H&E) histology would enable single-cell analysis at scale, but training such models has relied on manual pathologist annotations, which are slow, expensive and unreliable for many cell types. We instead supervise morphology with molecules. Imaging-based spatial transcriptomics profiles individual cells in situ on a section that can afterwards be stained with H&E, so that molecular identity and morphology are observed for the same physical cell. We assembled 81 such paired Xenium sections spanning 16 organs, derived per-cell labels by clustering, marker-gene annotation, organ-wise human review and quality control, and mapped them onto the cell types commonly reported in each organ. This yielded 15.4 million cells, each with a paired H&E image patch and one of 23 cell types, on which we trained CytoFormer, a cell foundation model with a multi-task, per-organ classification head. On spatially held-out tissue CytoFormer reached an accuracy of 0.85 and a macro-F1 of 0.78 across all 16 organs, and its predictions reproduced the tissue architecture of an entire held-out section. The representation also transfers: with the encoder frozen, a linear head on CytoFormer features performed better than six pathology foundation models on four expert-annotated benchmarks, including on organs and cell types that were not part of pretraining. Finally, in an interactive active-learning setting, CytoFormer's embeddings are markedly more label-efficient than existing pathology foundation models, detecting normal epithelium amid look-alike tumour with an F1 of 0.82 from only a few annotations and leading the strongest baseline by 0.13 in F1. CytoFormer turns paired H&E and spatial transcriptomics into a reusable, label-efficient representation for cell-level analysis of routine histology.

Jialu Yao, Songhao Li, Alina Yu. et al. · 0 citations
Open access Aug 2026

Cytoflow: User-Friendly Python Software for Computational Flow Cytometry.

Modern flow cytometry experiments routinely measure 18 or more fluorescent markers across many samples, patients, and tissues. As these experiments' complexity increases, manual gating becomes unacceptably inefficient and can introduce operator-to-operator variation. Computational cytometry algorithms such as unsupervised clustering, automated gating, and dimensionality reduction can increase the speed and reliability of these analyses, but using them requires time and expertise that many biomedical scientists lack. Cytoflow was built to bridge this gap. Cytoflow is open source, user-friendly point-and-click software written in Python that allows non-programmers to apply modern computational flow cytometry methods to their data sets. Additionally, Cytoflow's modules can be used directly in a Python script or a JupyterLab notebook. Finally, extending Cytoflow with new modules that support future applications is straightforward, and feedback from an active user community continues to guide ongoing development. As a result, Cytoflow can save an experimenter time and improve the reliability and reproducibility of their analysis. Cytoflow is available to download at https://cytoflow.github.io. Source code is hosted at https://github.com/cytoflow/cytoflow, and documentation is available at https://cytoflow.readthedocs.io.

Leah Teague · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.