Skip to content
Book Open access

Concord: Building Consensus Representations for Single Cells with Collaborative Random Projection

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · 0 citations · 47 references

TL;DR

Concord is a novel single-cell FM that explicitly models relationships between gene identities and expression levels through a collaborative rotary attention (CRA) mechanism, and employs two collaborative attention modes: a gene-to-expression rotational attention that produces gene-enhanced expression representations, and an expression-to-gene rotational attention that yields expression-enhanced gene representations.

Abstract

Foundation models (FMs) have recently transformed single-cell genomics by learning transferable representations from large-scale single-cell data, enabling a wide range of downstream biomedical applications. Inspired by natural language processing, existing single-cell FMs adapt transformer architectures by treating genes as tokens and cells as sequences. However, transformers are inherently agnostic to input order, while genes lack a natural sequential structure. Current approaches rely on heuristic strategies, such as expression-based gene sorting, to impose positional information, which often fail to capture relative relationships among collectively expressed genes and between gene identities and their expression levels, leading to information loss and limited generalization. In this work, we propose Concord, a novel single-cell FM that explicitly models relationships between gene identities and expression levels through a collaborative rotary attention (CRA) mechanism. Specifically, Concord employs two collaborative attention modes: a gene-to-expression rotational attention that produces gene-enhanced expression representations, and an expression-to-gene rotational attention that yields expression-enhanced gene representations. These two processes provide distinct yet complementary views of the same cell-level expression profile; accordingly, we employ contrastive learning to align their semantic representations. Our theoretical analysis demonstrates that CRA effectively leverages the relative distances in one embedding space to establish stable and consistent dependencies in the other modality. Extensive experiments on various single-cell datasets demonstrate that Concord outperforms existing FMs across various downstream tasks and provides more informative and transferable gene and cell representations. Our code is available at https://github.com/Catchxu/Concord.

Read PDF

Similar papers

Open access Sep 2026

scRep: A Latent-Space Self-Distilled Foundation Model for Single-Cell Representation Learning

Single-cell foundation models have shown strong potential for learning transferable representations from large-scale transcriptomic data. However, many existing approaches rely on reconstructing masked gene expression values, creating a potential mismatch between observation-space reconstruction and the goal of learnin...

Sheng-Jie Wang, Zong-Yong Hu, Yunlong Bie et al. · 0 citations
Sep 2026

scGFormer: A Multi-Scale Graph-Transformer for Cell Type Annotation in Single-Cell RNA Sequencing.

ScGFormer is equipped with a biology-guided adaptive contrastive learning strategy, which is designed to account for zero inflation, balance class distributions, and refine dynamic graphs during training, thereby facilitating robustness and adaptability.

Ziqi Yuan, Hong-Wei Zhang, Cheng Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Towards a knowledge-enhanced single-cell foundation model

Single-cell foundation models (scFMs) increasingly rely on large-scale transcriptomic pretraining, yet expanding pretraining data can yield diminishing gains while substantially increasing computational cost. Our data scaling analyses showed that incorporating biological knowledge, including cell-level text annotation...

Han-Qing Zhang, Jie Bao, Mei Ma et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CellMSA: Context Modeling for Single-Cell Representation Learning

Single-cell transcriptomics enables profiling of cellular states at unprecedented resolution, but its high dimensionality, sparsity, and technical batch effects pose significant challenges for representation learning. Existing single-cell foundation models typically encode each cell independently or only model cells fr...

Su-Yuan Zhao, Ming-Hao Liu, Yi-Zhen Luo et al. · 0 citations
Open access Sep 2026

RAGCell: Retrieval-Augmented Generation as Supervision for Versatile Single-cell Analysis.

MOTIVATION Single-cell foundation models (scFMs) are transforming computational biology by enabling generalizable, task-agnostic representations for versatile single-cell analysis. Despite their progress in facilitating rapid deployment for downstream tasks, off-the-shelf scFMs still have some overlooked concerns: (I)...

Tian-Yu Liu, Fan Zhang, Jia-Yuan Chen et al. · 0 citations
Preprint Aug 2026

bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning

This work proposes bioMoR, which is the first framework to apply MoR to gene-level and pathway-level learning, and identifies three locations for integrating structured biological knowledge within an MoR backbone: graph-based information sharing refines token embeddings, a structural bias guides self-attention toward b...

Koushik Howlader, Tirtho Roy, Md Tauhidul Islam et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.