Jul 2026· International Journal of Data Science and Analysis· Vol 22· 0 citations· 51 references
Computer Science
TL;DR
HiTOC is proposed, a Hierarchical Topological Ordering framework for Causal discovery that constructs the global causal structure via layer-wise integration of local orderings over induced subsets by Markov Blanket and achieves state-of-the-art performance in both accuracy and scalability, particularly in high-dimensional and structurally complex settings.
Abstract
Causal discovery aims to recover the underlying directed acyclic graph (DAG) from observational data. Ordering-based methods offer a scalable perspective to global DAG search by estimating a topological order and assigning edges accordingly. However, existing approaches depend on repeated full-graph score evaluations and assume that local score minima correspond to causal sinks, an assumption that breaks down under statistical noise, model misspecification, or dense local dependencies. To overcome these limitations, we propose HiTOC, a Hierarchical Topological Ordering framework for Causal discovery that constructs the global causal structure via layer-wise integration of local orderings over induced subsets by Markov Blanket. At each iteration, nodes identified as sinks through local score-based rankings are peeled off to form hierarchical layers, avoiding global permutation search. HiTOC provides a clearer hierarchical structure via recursive modular inference and sink extraction, enabling interpretable layer-wise inference and localizing potential errors to small subgraphs. To ensure robustness under local inconsistencies, we further introduce a calibration mechanism with theoretical guarantees. Empirical results demonstrate that HiTOC achieves state-of-the-art performance in both accuracy and scalability, particularly in high-dimensional and structurally complex settings.
Causal Discovery (CD) from observational data faces two fundamental challenges. First, purely statistical methods often lack the power to resolve structural ambiguities in low-sample regimes. Second, although LLM-assisted hybrid approaches improve structure recovery through semantic reasoning, the influence of that reasoning on individual edge decisions remains largely opaque. Consequently, existing hybrid methods fail to satisfy a fundamental requirement: explaining why a particular edge is included or excluded in the learned directed acyclic graph (DAG). This is critical in real-world applications, where no ground-truth DAG exists and every structural decision must be independently justified. We formalize this requirement as decision traceability, requiring every inferred edge to be supported by auditable statistical evidence, Markov Blanket consistency, or explicit domain reasoning. We propose GENESIS, an explainable hybrid CD framework that decomposes graph construction into interpretable decision points. GENESIS first identifies and scores three-node structural motifs, including chains, forks, and colliders, to establish transparent structural priors, then progressively refines the graph by integrating these priors with observational evidence, invoking domain knowledge only when statistical evidence is insufficient. By design, every edge decision is resolved through an auditable source of evidence. Experiments show that GENESIS achieves 100% decision traceability across all settings, establishing explainability as a first-class objective in causal discovery. Despite this additional requirement, GENESIS consistently outperforms purely statistical CD methods on the majority of benchmark datasets across all sample regimes in terms of Structural Hamming Distance (SHD), while achieving performance comparable to state-of-the-art LLM-assisted approaches.
A. Thorat, Ravi Kolla, Vishak K Bhat et al.· 0 citations
SVI-DAG is proposed, a structured variational inference approach to Bayesian causal discovery using observational data and prior beliefs that uses normalizing flows to model dependencies between edges, supporting expressive and multimodal posterior learning over DAGs.
The results show that representing both sources as probabilistic uncertainty over edge existence and orientation is a practical and effective way to improve causal graph accuracy.
N. K. Kitson, Anthony C. Constantinou· 0 citations
Clustering populations of networks while recovering their latent hierarchical organization is a fundamental yet largely unexplored problem in network analysis. To formalize this, we introduce the Hierarchical Distance Matrix, a specific class of population-level distance matrices that encodes latent hierarchical organization through recursively nested distance separation, accommodating unbalanced tree depths. Building on this framework, we propose a fully data-driven top-down procedure: network hierarchical clustering based on two-sample testing (NHC-TST). The algorithm recursively splits networks via spectral clustering and uses a graph-based two-sample stopping rule. The procedure adaptively determines the branching structure without requiring prior knowledge of the number of clusters or tree depth. Theoretically, we establish exact recovery of the population-level hierarchical structure and statistical consistency in the empirical procedure. Simulation studies demonstrate highly accurate recovery of both cluster memberships and hierarchical relationships across a wide range of settings. Applied to a global migration dataset, NHC-TST uncovers interpretable multi-resolution temporal structures that are not revealed by conventional flat clustering approaches.
Li Chen, Nathaniel Josephs, E. Kolaczyk et al.· 0 citations
This work considers the task of conditional causal discovery as a Bayesian inference problem, in which the posterior is targeted over causal graphs and parameters conditional on an event such as a causal-effect constraint, and adapts rare-event estimation techniques to perform inference the joint graph-parameter space.
Cixuan Zhang, Guy Van den Broeck, Benjie Wang· 0 citations
The Hierarchical Heterogeneous Information Network is proposed as a data model that organizes typed entity relations in a main structure layer and descriptive evidence in a strong attribute layer as well as developing controllable approximate reduction as one instantiation.
Qinggeng Jin, Wujie Hu, Yongjie Liang et al.· Journal of King Saud Univers...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.