Aug 2026· International Journal of Data Science and Analysis· Vol 22· 0 citations· 35 references
TL;DR
SaDGAE, an unsupervised deep graph autoencoder (GAE) framework that jointly models gene expression patterns and cell–cell relationships, achieves strong and competitive clustering performance, yielding biologically interpretable clusters and accurately recovering known marker gene patterns.
Single-cell RNA sequencing (scRNA-seq) provides a novel perspective to explore cellular biology at the single-cell resolution. Single-cell clustering is a crucial step to reveal cell types and the corresponding biological functions. However, when dealing with the high dimensionality and complexity of scRNA-seq data, existing deep models fail to comprehensively capture the intrinsic attribute information and structural relationships within the data. In this study, we propose a novel single-cell deep clustering model named scDFVA. The proposed scDFVA consists of a variational graph attention autoencoder (AE), a zero-inflated negative binomial (ZINB) based AE, and a self-supervised clustering. To better simulate sparse and zero-inflated scRNA-seq data, we incorporate the ZINB model into the AE. The variational graph attention AE is introduced to learn the cell structure information. scDFVA achieves representation learning within a joint framework comprising a ZINB-based AE and a variational graph attention AE, effectively fusing gene expression and cell structure information. Furthermore, scDFVA performs self-supervised clustering training on the latent fusion representations of cells to achieve mutual supervision between representation learning and clustering. Experiments indicated that scDFVA outperformed several other competing methods, demonstrating that our method is beneficial in single-cell clustering.
Ge Zhang, Maohua Qin, Xuye Kou et al.· J. Comput. Biol.· 0 citations
Single-cell RNA sequencing (scRNA-seq) enables transcriptomic profiling at single-cell resolution, but accurate identification of cell subpopulations remains challenging because of the high dimensionality, sparsity, and dropout effects of scRNA-seq data. Existing deep learning-based clustering methods have shown promising performance, yet many primarily emphasize local neighborhood aggregation and may fail to adequately capture long-range cellular dependencies. Here, we propose Synergistic Global Transformer and Adaptive Graph Gating for Accurate scRNA-seq Clustering (AGTformer), an unsupervised clustering framework that combines adaptive edge reweighting with global latent-space modeling. AGTformer employs an Adaptive Adjacency Gating mechanism to dynamically reweight existing edges in the initial cell-cell graph, thereby reducing the influence of unreliable local connections and improving the stability of topology-aware representation learning. It further incorporates a Global Transformer refinement module to model long-range cell-cell dependencies beyond local graph propagation. Through the synergy of local topology-aware learning and global contextual refinement, AGTformer learns discriminative latent representations for clustering. Experiments on ten public scRNA-seq datasets demonstrate that AGTformer achieves superior clustering performance over representative baseline methods. In addition, visualization, sensitivity analysis, and ablation study support the effectiveness of the proposed components in improving representation quality for scRNA-seq clustering. These results suggest that AGTformer is a useful framework for unsupervised characterization of cellular heterogeneity in single-cell transcriptomic data.
Yuanyuan Dang, Wenqiang Liu, Hao Li et al.· Computational biology and ch...· 0 citations
The framework integrates three modules: dual-reconstruction to fuse attribute-structure information, contrastive learning under label guidance to extract semantic similarities, and deep embedding clustering to enable iterative optimization.
Wenjing Su, Baojuan Qin, Junliang Shang et al.· Interdisciplinary Sciences C...· 0 citations
Experiments demonstrate that the GraphFusionNN fusion strategy— integrating graph topology, spatial context, and external embeddings—significantly improves classification accuracy, macro-F1, and robustness compared to single-modality models.
Data from single-cell RNA sequencing (scRNA-seq) and the Assay for Transposase-Accessible Chromatin (scATAC-seq) are high-dimensional, sparse, and undesirably capture technical variability between experiments or batches. Many analysis methods thus seek to produce a low-dimensional cell-by-feature embedding space that groups together biologically similar cells across batches while distancing dissimilar cells. Here, we introduce ensemble refinement for scRNA-seq and scATAC-seq embeddings, inspired by ensemble methods from statistical machine learning, and implement BatchRefiner, a fast post-processing tool to enhance batch integration. We extensively benchmark widely-used scRNA-seq embedding methods on both batch integration and biological conservation over a wide range of datasets, before and after the addition of BatchRefiner. We extend these benchmarking approaches to provide the first comprehensive benchmark of batch integration for scATAC-seq embedding methods, including BatchRefiner. Importantly, we formalize a significance statistic, which we use to demonstrate BatchRefiner’s significant improvement in batch integration across a wide range of embedding methods, atlas-scale datasets, and established metrics.
Daniel E. Schäffer, Helen Kang, E. D. Aksu et al.· bioRxiv· 0 citations