Skip to content

Mapping the DAG-ness Landscape: Structural Archetypes in Complex Networks

Jul 2026 · arXiv.org · Vol abs/2607.16464 · 0 citations · 47 references
Computer Science

TL;DR

This paper empirically evaluates the DAG-ness framework, a four-component measure that quantifies acyclicity, flow alignment, cyclic locality, and pathway complexity across a corpus of 107 networks drawn from twelve structurally diverse domains, and finds that macroscopic acyclicity is pervasive even in feedback-rich systems.

Abstract

Directed networks arise across biological, social, informational, and engineered systems, yet most analyses treat directedness as a binary property: a network is either a directed acyclic graph (DAG) or it is not. This binary classification obscures the rich spectrum of hierarchical, recurrent, and modular structure present in real systems. In this paper, we empirically evaluate the DAG-ness framework, a four-component measure that quantifies acyclicity, flow alignment, cyclic locality, and pathway complexity across a corpus of 107 networks drawn from twelve structurally diverse domains. Rather than aligning with traditional disciplinary boundaries, our results reveal unexpected cross-domain convergence: diverse systems resolve into four universal structural archetypes. We find that macroscopic acyclicity is pervasive even in feedback-rich systems, and that domains as disparate as neural connectomes and abstract informational networks frequently converge on identical topological constraints. These findings demonstrate that DAG-ness provides a unified, interpretable, and domain-agnostic lens for understanding the hidden laws of directed structure in complex systems.

View source

Similar papers

Jul 2026

ZipLine: Visual Analysis of Multivariate Graphs with Predicate Logic

Multivariate graphs unite two distinct data perspectives: a topological structure defined by nodes and edges, and attribute data associated with each node. Analyzing such graphs therefore requires reasoning across two complementary spaces. However, existing systems typically emphasize the analysis of one space at a time, focusing either on topology or on attributes. As a result, exploration, analysis, and pattern discovery that depend on their interaction remain difficult. In this paper, we present ZipLine, a system designed to support integrative analysis of multivariate graphs by bridging both topology and attribute spaces. ZipLine introduces a predicate language that enables analysts to express patterns involving topology, node attributes, and neighborhood relations with a unified formalism. The system further provides a predicate-learning algorithm that maps analyst interactions across both topology (e.g., subgraph selection) and attribute views (e.g., value brushing), into the predicate language, enabling learned expressions that bridge the two spaces. This integrative approach supports iterative analysis by enabling analysts to refine patterns through coordinated reasoning over topology and attributes. We demonstrate ZipLine through three case studies in energy infrastructure, cybersecurity, and drug discovery analysis. The results show that ZipLine enables expressive multivariate graph analysis through unified reasoning across topology and attributes.

Sjoerd Vink, Suyang Li, Brian Montambault et al. · 0 citations
#machine learning Preprint Sep 2026

Language-encoded network topology enables large language models to reason about complex networks

Networks describe systems in biology and beyond, from protein interactions and social relationships to power grids and citation records. Reasoning about such systems requires understanding their structure: which elements are central, which connections bridge separate communities, and how it changes when elements are removed. Although large language models (LLMs) excel at natural language, they struggle with such questions when networks are given as edge lists, sentences or measurement tables, because their structural meaning must be inferred. Here we introduce BioGlyph, which compiles network topology into an interpretable and transferable language of structural roles. BioGlyph combines graph partitioning and structural measurements to identify roles such as hubs, community cores and cross-community connectors, and fixed rules to translate them into a universal vocabulary. The representation describes each element through its structural role, supporting evidence and semantic consequences, leaving both the network and the LLM unchanged. Across twenty networks spanning five domains, BioGlyph substantially improves open LLMs'ability to answer structural reasoning questions, outperforming edge-based, numerical and learned representations by up to 26 percentage points in system accuracy. Ablations show that the gain comes from explicitly encoding structural roles in semantically interpretable terms. The gain is more prominent in dense, community-structured networks and diminishes in sparse networks whose topology is more readily inferred from text. In a budding-yeast protein-interaction network, BioGlyph exposes biological organization: cross-community connectors are enriched for essential genes, whereas peripheral proteins are depleted. BioGlyph thus provides an interpretable representation for both language models and scientists to reason about network structure.

Ucchwas Talukder Utsha, Sakib Mostafa, James Zou et al. · 0 citations
Jul 2026

Link Prediction Based on Subgraph Learning in Biological Networks

Recently, link prediction (LP) based on graph neural networks (GNNs) methods has achieved notable successes in biological networks (BNs), since it can reveal the organizational principles, functional mechanisms, and dynamic properties of biological systems. However, these LPs still face some significant challenges that need to be addressed in BNs. The first is the complex and heterogeneous characteristics of BNs. Moreover, BN structures often have the dynamic addition and removal of nodes and edges over time or across physiological states. Afterward, there are asymmetries and hierarchical modularity in the structures of BNs. Finally, high computational complexity has resulted from the above challenges in BNs. Therefore, this article proposes a novel GNN-based LP model via local clustering and subgraphs, termed LCS in short, to address these issues in BNs. LCS introduces a subgraph-based GNN approach to effectively address the heterogeneous characteristics inherent in complex BNs, along with the consequent challenges of asymmetry and hierarchical modularity. Furthermore, LCS designs a dynamic local subgraph extraction (SE) mechanism based on heat kernel diffusion and the Chopper pruning algorithm. This mechanism leverages the effective local clustering properties of heat diffusion and uses Chopper to achieve linear-time SE, mitigating subgraph size explosion and enhancing LP efficiency. Additionally, by imposing diversity regularization constraints, the method reduces computational complexity and improves generalization performance. Experimental results on four BN benchmarks demonstrate that LCS achieves significant improvements over existing state-of-the-art LP methods. The implementation of LCS is publicly available at https://github.com/XL0104/LCS-Model.git.

Xiao-Long Liu, Jianxia Chen, Wenzhe Chen et al. · 0 citations
Review Aug 2026

Topology as a language for emergent organization in complex systems: Multiscale structure, higher-order interactions, and structural diagnostics.

Complex systems are difficult to study not only because they are nonlinear, multiscale, and nonstationary, but because their scientifically relevant organization is often distributed across components, relations, and interaction orders. Topology provides a mathematical language for describing that organization through connectedness, recurrence, branching, closure, cavities, and persistence across scale. This review synthesizes persistent homology, Mapper, simplicial complexes, hypergraphs, and relation-level operator methods through a unified workflow from empirical data to representation, topological construction, output, and scientific interpretation. Across nonlinear dynamics, finance, neuroscience, biology, ecology, materials, and engineered systems, topological and topology-inspired methods make state-space organization, collective constraints, and structural reorganization available as observables that can be integrated with statistics, dynamics, mechanistic models, and machine learning. The review distinguishes the claim that a representation makes structure visible from the stronger claim that it improves detection or prediction, and it summarizes comparative evidence where such benchmarks exist. Prospective early-warning evidence remains uneven, but several studies demonstrate useful structural diagnostics, data-efficient classification, anomaly detection, and reductions in false alarms. The central conclusion is that topology is most valuable when representation is treated as a scientific hypothesis and topological descriptions are connected to domain-matched inference and mechanism.

Mark M. Bailey · 0 citations
Preprint Aug 2026

Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion

This work proposes Graph Diffusion Counterfactual Explanation via Inversion (GDCE-I), a discrete denoising diffusion model with a novel discrete inversion scheme that enables distribution-aware edits leveraging the whole domain edit space and qualitatively shows that GDCE-I attains interpretable in-distribution solutions.

David Bechtoldt, Sidney Bender · 0 citations
Review Aug 2026

A Graph Approach to the Academic Publishing Network: A Heterogeneous Model and Structural Screening over OpenAlex Open Data

The academic publishing ecosystem is a vast, heterogeneous network of works, authors, institutions, journals, and topics. Traditional scientometrics reduces it to isolated tabular indicators (h-index, Impact Factor) that ignore topological context and are not designed to capture coordinated illegitimate practices. Building on our companion review, which proposed graph analysis of publishing integrity, this paper implements that approach. We define a heterogeneous multivariate graph model over OpenAlex open data (seven node types, seven edge types) and a methodology based on projections (citation and co-authorship networks), interpretable structural metrics, community detection, and three screening detectors of anomalous publishing patterns. We deliberately avoid binary classification: detectors return ranked candidates with explicit structural evidence for human assessment. On the institutional corpus of VSB - Technical University of Ostrava (2020-2025) with its one-hop citation neighbourhood, community detection recovers real research groups, centralities identify cross-disciplinary bridges, and the screenings flag dense co-authorship cliques, locally closed citation loops, and thematically isolated venues. On a second, venue-centric corpus with external ground truth (journals delisted by Scopus and DOAJ) and size-matched controls, a naive case-control design yields seemingly strong but spurious detectors (a prominence confound), whereas after matching the only robust signal is the breadth of disciplinary scope (AUC 0.70); an open graph-based prestige measure (PageRank over the journal citation network) tracks a JIF proxy while being an order of magnitude more resistant to citation gaming than count-based indicators. We release the method as the open-source library apnet with a reproducible CLI workflow and a web interface; the analysis runs on commodity hardware in minutes.

Robert Šamárek, R. Martínek · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.