scRep: A Latent-Space Self-Distilled Foundation Model for Single-Cell Representation Learning
Abstract
Single-cell foundation models have shown strong potential for learning transferable representations from large-scale transcriptomic data. However, many existing approaches rely on reconstructing masked gene expression values, creating a potential mismatch between observation-space reconstruction and the goal of learning stable biological representations. This challenge is particularly relevant to single-cell RNA sequencing, where sparsity, incomplete gene detection, and technical variation can obscure the underlying biological state. Here, we introduce scRep, a compact latent-space self-distillation framework for single-cell representation learning. Rather than reconstructing raw expression values, scRep aligns differently perturbed views of the same cell through a momentum-updated teacher–student architecture, with self-distillation objectives at both the cell and gene levels. This representation-centered formulation encourages the model to capture biological information that remains stable across incomplete and perturbed transcriptomic observations. Using frozen representations without task-specific fine-tuning, scRep pretrained on approximately 2.8 million cells achieves the strongest overall performance across the evaluated frozen-representation benchmarks, demonstrating strong sample efficiency. A larger-scale scRep model pretrained on 30.72 million cells further demonstrates that the framework remains effective when scaled to a substantially larger and more diverse corpus. Beyond cell identity, scRep prioritizes established marker genes, recovers transcription factor–associated gene programs with cell-type-specific activity, and preserves continuous developmental structure that supports graph-based pseudotime inference. We further show that pretraining performance is closely associated with biological diversity: reducing redundant cells while improving cell-type coverage can match or exceed the performance of larger, less balanced training corpora. Together, these results establish latent-space self-distillation as an effective alternative to expression reconstruction for single-cell foundation modeling and suggest that efficient scaling depends not only on the number of cells, but also on the learning objective and the biological diversity of the pretraining corpus.