The results indicate that GmGM provides a unified, reproducible framework for joint cell clustering and gene-network inference, capable of revealing cellular structure beyond that captured by conventional pipelines.
Abstract
Conventional annotation of single-cell RNA-sequencing (scRNA-seq) data relies heavily on manual, marker-based thresholding, an approach that can obscure subtle transcriptomic gradients and collapse functionally distinct cell states into broad, heterogeneous populations. Here we apply the Gaussian multi-Graphical Model (GmGM) framework, which jointly infers cell-cell and gene-gene dependency structure from a single scRNA-seq data matrix, to a 10x Genomics PBMC dataset. Ten independent GMGM-Leiden clustering runs were integrated into a robust consensus partition using a soft cluster ensemble approach and benchmarked against reference cell-type annotations. This strategy yielded stable cluster partitions that resolve biologically meaningful sub-populations not distinguished by the reference annotation. In parallel, for each cluster, gene co-expression modules were extracted from the fitted model via consensus Leiden clustering across resolutions, evaluated using standard network metrics, and validated functionally with the Network Enrichment Analysis Test (NEAT), which confirmed non-random enrichment signal. A module-scoring procedure linked network topology to per-cell, per-cluster expression signatures, and a novel extension of GmGM, recovering a shared cell-cell network together with population-specific gene networks in a single model run, was demonstrated in a case study on the CD4+ T-cell population. These results indicate that GmGM provides a unified, reproducible framework for joint cell clustering and gene-network inference, capable of revealing cellular structure beyond that captured by conventional pipelines.
Feature selection is critical for resolving cell-type heterogeneity in single-cell RNA sequencing (scRNA-seq). DUBStepR (Determining the Underlying Basis using Stepwise Regression) is a widely used gene selection method for scRNA-seq designed to identify feature genes that maximize cell-type separation. DUBStepR has been reported to perform effectively in this domain; however, its reliance on linear Pearson correlation and rigid thresholding limits its effectiveness on complex, high-dimensional datasets. Three enhancements are presented in this study: RFCell-DUBStepR, which uses random forests to capture expression-level importance; Copula-DUBStepR, which models non-linear correlations via Gaussian Copulas; and Zqt-DUBStepR, which utilizes quantile-based selection for improved gene retention. Using both simulated and real-world datasets (scRNA-seq), these modifications are shown to resolve the biases of the original algorithm. The modified methods consistently select a more representative gene set and yield higher clustering accuracy across varying levels of biological complexity. These findings establish the modified DUBStepR frameworks as more reliable tools for high-fidelity subpopulation identification in downstream single-cell analysis.
Single-cell RNA sequencing (scRNA-seq) has opened unprecedented possibilities to explore the complexity of the immune system. However, existing methods primarily rely on expression-based clustering analysis, which lacks mechanistic explanations for immune cell states and encounters challenges in integrating multi-scale data.
We developed a network-informed deconvolution framework that constructs Bayesian network-derived regulatory structures using immune-related genes from context-matched bulk RNA-seq datasets. Network markers were extracted from these structures and projected onto peripheral blood mononuclear cell (PBMC) and lung adenocarcinoma (LUAD) scRNA-seq datasets to identify network biomarkers and define immune cell states. Spatial transcriptomic analysis was further used to evaluate the spatial coherence of network-defined cell states. The scRNA-seq and spatial transcriptomic datasets analyzed in this study were generated from prospectively collected samples by our team, while context-matched bulk RNA-seq cohorts were used to derive population-level immune gene network structures.
The framework identified structure-defined immune subpopulations in both PBMC and LUAD datasets and revealed functional heterogeneity across multiple immune lineages. Spatial transcriptomic analysis further showed that network-associated immune clusters exhibited closer spatial proximity than non-associated clusters, supporting the spatial coherence of network-defined cell states.
This framework provides a network-informed representation for immune cell subpopulation identification and functional characterization. By linking bulk immune gene co-expression, single-cell programs, and spatial organization, this approach offers an additional perspective for understanding immune dynamics in both normal and pathological states and may provide an analytical basis for more precise immunotherapy-related studies.
Yi-Ming Li, Yujie You, Ruixian Chen et al.· Frontiers in Immunology· 0 citations
MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer–based statistics to detect biological variation without requiring alignment, is presented.
L. Mboning, Maciej Dlugosz, Marek Kokot et al.· bioRxiv· 0 citations
INTRODUCTION
Single-cell RNA sequencing (scRNA-seq) data exhibit extreme sparsity, technical noise, and complex nonlinear structures that obstruct accurate biological interpretation. Current methods inadequately model intercellular relationships and suffer from feature selection instability.
PURPOSE
(1) Automatically identify biologically relevant features in noisy scRNA-seq data,(2) Model multi-scale cellular relationships for robust clustering,(3) Provide interpretable representations for mechanistic insights.
METHODOLOGY
Adaptive HVG selection: Improved Random Forest to mitigate dropout artifacts. Hybrid graph autoencoder: Fusion of Graph Attention Networks (local interactions) and Graph Convolutional Networks (global neighborhoods). Biologically informed optimization: MMD regularization + Spearman-correlation feature filtering.
RESULTS
Evaluated across 17 scRNA-seq datasets: Showed better clustering performance than CellVGAE in our experiments (Silhouette Coefficient and Davies-Bouldin Index),Selection of highly variable genes mitigated technical noise while enhancing biological signal retention.
CONCLUSION
GCAN establishes a new paradigm for scRNA-seq analysis by unifying adaptive feature selection with context-aware relational modeling. Its architecture implements adaptive screening of biologically relevant features and enables accurate cell typing.
Jingyu Bai, Li Xu· Computational biology and ch...· 0 citations
Single-cell RNA sequencing technology dramatically changed the way we investigate transcriptomes. However, the amount and complexity of data generated by such methods poses new challenges for biologists who are trying to extract detailed insights into the genetic programs that drive cellular functions and differentiation. To provide a more intuitive understanding of cell specific gene expression programs, we developed a novel approach for exploiting scRNA-seq data that detects individual gene expression levels in each cell, by avoiding dimensional reduction methods. This was achieved by focusing our analysis on individual cells with a high sequencing coverage (above 15000 Unique Molecular Identifiers (UMIs)). Such High Coverage Cells (HCC), were found in all five C. elegans scRNA-seq datasets we investigated and constitute direct quantitative experimental observations of the mRNA content of individual cells. Clustering the complete gene expression matrix for these cells, we identified gene sets specific for most C. elegans tissues. Among each set we found genes that are dominating cell specific transcriptomes as well as genes that are restricted to particular cell types but are a thousand fold less expressed. For each cell type or subtype we characterized, we identified a set of genes with expression restricted to those cells that were not previously associated with the corresponding tissue. Our results demonstrate that by focusing on HCCs, we can provide high-resolution quantitative descriptions of cellular expression landscapes that are immediately exploitable for researchers to generate new biological hypotheses. Overall, we demonstrate that HCCs represent a powerful and largely unexplored source of biological insights and suggest that future scRNA-seq experiments could benefit from focusing on HCC enrichment to capture and exploit the full complexity of cellular transcriptomes.
Florian Bernard, Emma Kandel, D. Dargère et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.