Gene regulatory networks (GRNs) are graph-based representations of regulatory relationships between transcription factors and target genes. Various tools exist to infer GRNs from gene expression data, but since this task is computationally intensive, statistical significance estimates are often omitted. While permutation-based empirical P-value computation methods are relatively straightforward to implement, they are prohibitively expensive when applied to popular regression-based GRN inference methods and realistically sized datasets. To address this bottleneck, we developed SignifiKANTE. SignifiKANTE is based on the key insight that the background count distributions of groups of target genes may be highly similar, even if their expression vectors show distinct behavior. Relying on this insight, SignifiKANTE employs gene clustering based on the 1-Wasserstein distance to create a small, constant number of background distributions which enables the simultaneous computation of approximate empirical P-values for multiple target genes. This reduces runtime by orders of magnitudes (for some datasets, from several weeks to few hours), without compromising faithfulness of the obtained P-values. SignifiKANTE extends the popular GRN inference package Arboreto and is available as a Python package on GitHub (https://github.com/bionetslab/SignifiKANTE) and PyPI (https://pypi.org/project/signifikante/).
F. Woller, Paul Martini, Souptik Sen et al.· bioRxiv· 1 citation
Motivation Graph Neural Networks (GNNs) have gained increasing interest in the biomedical domain, as the integration of prior knowledge and deep neural networks has the potential to enhance insights into molecular processes and disease mechanisms. However, a comprehensive and systematic assessment of model architectures, data modalities, graph structures, and their performance for graph signal classification in the biomedical domain is yet to be performed. In order to close this gap, we conducted a benchmarking study on multiple GNNs on a Protein-Protein Interaction (PPI) network for Kidney Renal Clear Cell Carcinoma and Breast cancer subtype prediction, performing an in-depth investigation of architectures, incorporating skip connections and various data modalities. Results While none of the GNNs outperforms the structure-agnostic Multi-Layer Perceptron baseline, all of them can handle bimodal data (gene methylation and expression) and offer the ability to gain explainability based on PPIs. We offer practical guidelines for applying GNNs to graph signal processing tasks specifically for cancer classification. Depending on the underlying dataset and PPI structure employed, models on different data modalities outperform others. Overall, we suggest using ChebNet, which tends to outperform the Graph Convolutional Network and the Graph Attention Network in cancer subtype prediction. We recommend using GNN architectures that employ a simple flattening readout layer, as they provide better classification performance and faster training time than those with global average pooling. Additionally, we tested residual connections, but they had only an insignificant impact on classification performance. Availability and implementation Code and data available at https://github.com/HauschildLab/GNN4PPI. Contact julia.schirmacher@med.uni-goettingen.de, anne-christin.hauschild@uni-gießen.de Supplementary information Supplementary data attached.
Julia Schirmacher, M. C. Maurer, J. M. Metsch et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.