Skip to content

LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and Orientation

Aug 2026 · 0 citations · 17 references
Computer Science

TL;DR

The results show that representing both sources as probabilistic uncertainty over edge existence and orientation is a practical and effective way to improve causal graph accuracy.

Abstract

Bayesian network structure learning (BNSL) from observational data struggles with orientation identifiability, while large language models (LLMs) offer broad but often unreliable causal knowledge. We propose combining these complementary sources through a novel representation, termed Probabilistic Dependency Graphs (PDGs). In a PDG, each edge is associated with a distribution over directed, undirected, and absent states, enabling fusion via weighted averaging. We evaluate this approach on 26 benchmark networks, combining ensembles of three BNSL algorithms (FGES, Tabu, PC) with three LLMs (Gemini, Claude, GPT) across multiple prompts and random seeds. A simple 50/50 fusion improves F1 over the better of either source alone in 22 of 26 networks, with a statistically significant mean improvement of $0.056$ $(p<0.001)$. Analysis reveals that the two sources play complementary roles: BNSL contributes a high-recall edge skeleton (80\% vs 60\% for LLM), while LLM contributes accurate edge orientation (96\% vs 77\% for BNSL). Our results show that representing both sources as probabilistic uncertainty over edge existence and orientation is a practical and effective way to improve causal graph accuracy.

View source

Similar papers

Preprint Aug 2026

SVI-DAG: A Structured Variational Inference Approach to Bayesian Causal Discovery

SVI-DAG is proposed, a structured variational inference approach to Bayesian causal discovery using observational data and prior beliefs that uses normalizing flows to model dependencies between edges, supporting expressive and multimodal posterior learning over DAGs.

Shrenik Zinage · 0 citations
Preprint Aug 2026

GENESIS: Towards Explainable Causal Discovery

Causal Discovery (CD) from observational data faces two fundamental challenges. First, purely statistical methods often lack the power to resolve structural ambiguities in low-sample regimes. Second, although LLM-assisted hybrid approaches improve structure recovery through semantic reasoning, the influence of that reasoning on individual edge decisions remains largely opaque. Consequently, existing hybrid methods fail to satisfy a fundamental requirement: explaining why a particular edge is included or excluded in the learned directed acyclic graph (DAG). This is critical in real-world applications, where no ground-truth DAG exists and every structural decision must be independently justified. We formalize this requirement as decision traceability, requiring every inferred edge to be supported by auditable statistical evidence, Markov Blanket consistency, or explicit domain reasoning. We propose GENESIS, an explainable hybrid CD framework that decomposes graph construction into interpretable decision points. GENESIS first identifies and scores three-node structural motifs, including chains, forks, and colliders, to establish transparent structural priors, then progressively refines the graph by integrating these priors with observational evidence, invoking domain knowledge only when statistical evidence is insufficient. By design, every edge decision is resolved through an auditable source of evidence. Experiments show that GENESIS achieves 100% decision traceability across all settings, establishing explainability as a first-class objective in causal discovery. Despite this additional requirement, GENESIS consistently outperforms purely statistical CD methods on the majority of benchmark datasets across all sample regimes in terms of Structural Hamming Distance (SHD), while achieving performance comparable to state-of-the-art LLM-assisted approaches.

A. Thorat, Ravi Kolla, Vishak K Bhat et al. · 0 citations
Preprint Aug 2026

From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers

Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, yet whether their direct-edge judgments and confidence can be trusted remains unclear. We systematically evaluate 12 instruction-tuned open-weight models across six benchmark causal graphs, five prompting strategies, and four confidence sources: verbalized, logit-based, cross-prompt agreement, and cross-model agreement. Under our language-only pairwise protocol, our evaluation yields three key findings. (i) LLM-based causal judgments are strongly recall-dominant: models predict overly dense graphs with many false-positive edges, while prompting mainly shifts the precision-recall trade-off rather than resolving overprediction. Gains from model scale diminish on the largest graphs and do not eliminate miscalibration. (ii) LLMs often capture causal relatedness without reliably identifying directness or orientation. Relative to published reference graphs, models misclassify 40.0% of indirect and 36.0% of reversed non-edges as direct edges, versus 28.2% of other non-edges. Moreover, 80.8% and 84.6% of these false positives receive verbalized confidence of at least 80%, revealing substantial overconfidence in structurally incorrect predictions. (iii) Conventional confidence estimates are unreliable, whereas agreement offers a more promising signal. Logit-based confidence frequently collapses near 1.0 regardless of correctness, while cross-prompt and cross-model agreement achieve better mean calibration and discrimination, though their advantages are not statistically significant after Holm correction. A benchmark-familiarity audit further identifies potential familiarity in five model-dataset pairs, all involving AsiaM. Overall, our results suggest LLMs are better viewed as sources of externally validated soft causal priors than as direct evidence of causal structure.

Amit Kumar, Elnur Adl Zarabi, S. Trivedy et al. · 0 citations
Open access Sep 2026

Advancing multilevel Bayesian networks with efficient Bayesian inference

Bayesian networks (BNs) are widely used tools for the modeling of complex dependencies among variables. However, their application to multilevel or clustered data remains constrained, especially in scenarios involving multiple random effects, due to a lack of methods and more pressing, efficient computational frameworks. In this article, we propose an innovative framework, INLA-MBN, that integrates multilevel BNs (MBNs) with the integrated nested Laplace approximation (INLA), thereby facilitating efficient structure and parameter learning in both longitudinal and cross-sectional multilevel data contexts. To the best of our knowledge, this is the first implementation of INLA for BNs that involves more than one correlated random effect. More precisely, we advance existing research by modeling longitudinal data with two correlated random effects and multilevel cross-sectional data with more than two correlated random effects. Our approach encompasses both Gaussian and hybrid MBNs that incorporate Gaussian and categorical nodes. Through an exhaustive simulation study, we demonstrate that INLA-MBN exhibit superior performance compared to networks based on maximum likelihood estimation, particularly, in scenarios characterized by small group sizes or limited repeated measurements per subject. INLA-MBN achieves a fully inferred MBN, while negating the cost associated with MBNs based on Markov Chain Monte Carlo. The proposed method was also applied to real-life longitudinal child morbidity data to demonstrate its practical applicability.

B. E. Yirdaw, L. K. Debusho, J. van Niekerk et al. · 0 citations
Open access Jul 2026

Hierarchical causal structure discovery from layered topological orderings

HiTOC is proposed, a Hierarchical Topological Ordering framework for Causal discovery that constructs the global causal structure via layer-wise integration of local orderings over induced subsets by Markov Blanket and achieves state-of-the-art performance in both accuracy and scalability, particularly in high-dimensional and structurally complex settings.

Haixiang Sun, Pengchao Tian, Zihan Zhou et al. · 0 citations
#machine learning Preprint Aug 2026

Causal Local States: Scalable Simultaneous Causal Network Inference and Forecasting for Dynamical Systems

Reconstruction of the underlying networks with high fidelity and forecasts on par with a model that is supplied with the true network are achieved, providing a step toward explainable and scalable forecasting of complex systems.

J. Braun, Fabian Fischbach, Daniel Köglmayr et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.