Skip to content
Preprint

Statistical Analysis of Network Collections Using Persistent Homology and Functional Data Analysis

Jul 2026 · 0 citations · 47 references
Mathematics

TL;DR

A framework termed functional topological data analysis (funTDA), which integrates tools from functional data analysis and topological data analysis to facilitate exploratory data analysis and inference on samples of networks, is introduced.

Abstract

Statistical analysis of collections of networks, where each network is treated as the primary unit of observation, is of growing importance across a wide range of application domains, including gene regulatory, social, and financial networks. As networks consist of vertices and edges that do not naturally reside in Euclidean space, the direct application of conventional statistical methodologies, such as the computation of means and covariances, principal component analysis, and hypothesis testing, to samples of networks is not straightforward. A central challenge lies in defining meaningful measures of similarity or distance between networks of potentially varying sizes and structural types (e.g., directed, undirected, weighted or unweighted), particularly when no predefined node correspondence exists. To address these challenges, we introduce a framework termed functional topological data analysis (funTDA), which integrates tools from functional data analysis and topological data analysis to facilitate exploratory data analysis and inference on samples of networks. The proposed framework enables the computation of summary statistics, including means and variances, and supports the application of principal component analysis and hypothesis testing to topological features extracted from network data. Through simulation studies involving networks with varying connectivity structures, we demonstrate the ability of funTDA to distinguish between distinct network configurations. The methodology is illustrated through two real-data applications: networks constructed from pairwise word co-occurrences in novels by Jane Austen and Charles Dickens, and gene regulatory networks derived from gene expression measurements for seventeen individuals exposed to H3N2 influenza. In both applications, differences in network topology are assessed using principal component analysis and hypothesis testing.

View source

Similar papers

Open access Jul 2026

A novel multivariate framework for functional gene networks enrichment analysis

Gene network analysis is critically implicated in disease research for uncovering functional modules and interaction-driven pathways underlying biological and disease processes. However, the interpretation of large inferred networks remains challenging. Although functional gene network analysis allows the interpretation of large inferred networks, challenges such as reduction of multiple network-level features to a single composite score often limit their application. This data reduction can mask the important multivariate characteristics of gene networks, hindering efficient differentiation of individual contributions of distinct network components. Hence, this study aimed to investigate a novel computational strategy called Multivariate Framework for Functional Gene Network Enrichment Analysis (mFGNA). This framework incorporated diverse network-level features from a graph-theoretical perspective, including node properties (centrality), edge connectivity patterns (Jaccard distance), interaction strengths (edge weights), alongside traditional expression levels. Notably, mFGNA preserved these multidimensional characteristics, capturing complex rewiring of gene networks across different phenotypic states. Furthermore, mFGNA adopted a gene-level permutation strategy to evaluate the enrichment hypothesis, ensuring effective statistical inference and reduced computational complexity compared with phenotype-based permutations. Extensive Monte Carlo simulations validated mFGNA through both undirected and directed gene networks, showing consistently improved performance over existing approaches across diverse pathway settings. We also applied mFGNA to investigate immune pathway perturbations in cancer cell lines and identified significant network-level dysregulation in pancreatic and non-small cell lung cancers. Cancer-specific interaction modules were dominated by human leukocyte antigen class II genes. Meanwhile, normal cell networks were characterized by hub genes such as MMP1 and MMP3 that were implicated in tissue maintenance, highlighting immune remodeling in tumors and the potential molecular targets for developing diagnostic and therapeutic interventions. Overall, the study shows that mFGNA enables effective functional pathway discovery in complex gene networks, providing mechanistic insights and potential translational targets in disease contexts.

Heewon Park, S. Imoto · 0 citations
Open access Nov 2025

SignifiKANTE: efficient P-value computation for gene regulatory networks

Gene regulatory networks (GRNs) are graph-based representations of regulatory relationships between transcription factors and target genes. Various tools exist to infer GRNs from gene expression data, but since this task is computationally intensive, statistical significance estimates are often omitted. While permutation-based empirical P-value computation methods are relatively straightforward to implement, they are prohibitively expensive when applied to popular regression-based GRN inference methods and realistically sized datasets. To address this bottleneck, we developed SignifiKANTE. SignifiKANTE is based on the key insight that the background count distributions of groups of target genes may be highly similar, even if their expression vectors show distinct behavior. Relying on this insight, SignifiKANTE employs gene clustering based on the 1-Wasserstein distance to create a small, constant number of background distributions which enables the simultaneous computation of approximate empirical P-values for multiple target genes. This reduces runtime by orders of magnitudes (for some datasets, from several weeks to few hours), without compromising faithfulness of the obtained P-values. SignifiKANTE extends the popular GRN inference package Arboreto and is available as a Python package on GitHub (https://github.com/bionetslab/SignifiKANTE) and PyPI (https://pypi.org/project/signifikante/).

F. Woller, Paul Martini, Souptik Sen et al. · 1 citation
#protein folding Review Open access Sep 2026

Graph-based representations in modern protein science

Abstract For more than 50 years, the linear sequence and the multiple sequence alignment have been the foundational data structures of protein science, and they remain central to homology search, phylogenetic inference, covariance-based contact prediction, and modern protein language models. However, relational and graph-based representations are increasingly being adopted alongside sequence-based methods to capture biological relationships that linear data structures express only implicitly. Proteins fold as three-dimensional residue interaction networks, evolve through high-dimensional genotype networks defined by mutational connectivity, and operate within cellular protein–protein interaction graphs. Here, we review how graph theory is being used to describe and understand these relationships across protein science, with an emphasis on what these methods offer biochemists working on enzyme superfamilies, protein engineering, drug targets, and functional annotation. We trace the development of these ideas from early theoretical topologies, through statistical coupling and the structural network analyses, to the geometric and graph-like representations used in recent machine-learning-driven advances. Throughout, we emphasise that graphs do not replace sequences or MSAs but provide a complementary representation for biochemical relationships that are difficult to express in one dimension.

Dana S. Matthews, Sacha B. Pulsford, Anthony Brancewicz et al. · 0 citations
Review Aug 2026

Network‐Valued Random Vectors and Their Statistical Foundations: A Comparison of Inferential Techniques and Testing Frameworks Under Heterogeneity

Network‐valued random vectors (NVRVs) provide a statistical framework for settings in which each observational unit is a network rather than a scalar, vector, image, or functional observation. Such data occur in social networks, omics and gene‐regulatory systems, functional brain connectivity, policy and intervention networks, and other domains where relational structure is itself the object of inference. NVRVs incorporate dependence through structured interactions among nodes and edges, thereby presenting significant challenges for statistical modeling, regression, and inference. This article provides an advanced review of inference techniques for NVRVs, with technical expositions focusing on the problem of measuring and testing network change in dynamic and heterogeneous settings. We review approaches for modeling networks as both responses and covariates, emphasizing key strategies such as edge‐wise models, summary‐based regression, latent variable methods, and Bayesian hierarchical formulations. We discuss how heterogeneity and temporal dynamics complicate inference, particularly in the context of network change detection, where current approaches frequently target isolated network features instead of the full network distribution. We consider the perspectives from both frequentist and Bayesian paradigms to identify fundamental gaps in current methodology including limitations in global network testing and the lack of theoretical guarantees in heterogeneous network data. Using the MRN‐114 dataset from the Mind Research Network database as empirical test cases, we perform integrative experiments demonstrating use of the techniques, discussing the advantages and pitfalls of each. Overall, our reviews compare techniques covering the representation of network‐valued observations, the measurement of dynamic change, and the inferential consequence of heterogeneity.

A. Roy, P. Auer, Anjishnu Banerjee · 0 citations
Preprint Aug 2026

muxvizpy: a Python library for the analysis of multilayer biological networks

Biological systems are inherently multilayered: the same entities---genes, cells, or bacterial species---participate simultaneously in qualitatively distinct types of interactions, each carrying complementary information that no single relational view can capture. Analysing such systems with single-layer tools, or by collapsing layers into a monoplex projection, systematically discards inter-layer dependencies and can yield misleading conclusions about centrality, community structure, and robustness. The multilayer network formalism addresses, and \texttt{muxViz} established one of the first comprehensive toolkits for its structural analysis, but its R-only interface and dense data structures limit applicability to large biological networks. We introduce \textit{muxvizpy}, a Python library that reimplements and extends the \texttt{muxViz} analytical catalogue with a sparse linear-algebra stack built on SciPy and PyTorch. Muxvizpy exposes seven categories through a unified, composable API and is numerically validated against \texttt{muxViz} on synthetic Erd\H{o}s--R\'enyi and Barab\'asi--Albert multiplex networks while substantially reducing peak memory and wall-clock time at scale. We illustrate its applicability on a virus--human protein-interaction multiplex in which computing some structural analysis was unfeasible. \\[2pt] muxvizpy is freely available under the MIT licence at https://github.com/CoMuNeLab/MuxVizPy. Mathematical definitions of all implemented metrics are provided in the Additional File.

Matteo Baldan, Francesco Zambelli, Andrea Valsecchi et al. · 0 citations
Review Open access Aug 2026

Graph Analytics for Social Networks

This paper conducts a case study on Zachary's Karate Club network and a synthetically generated scale-free network, computing centrality measures, detecting communities using the Louvain algorithm, and analysing degree-distribution behavior.

S. Sharma · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.