Skip to content

Local Causal Structure Learning in the Presence of Latent Variables and Selection Bias

Jul 2026 · arXiv.org · Vol abs/2607.19866 · 1 citation
Computer Science

TL;DR

LoCaLS is proposed, a local causal structure learning algorithm that is sound and complete under standard assumptions and identifies the same direct causes and effects of a target variable as those identifiable by global causal discovery methods, while allowing for latent variables and selection bias.

Abstract

Discovering the direct causes and effects of a target variable from observational data is a fundamental problem in causal discovery, with broad applications in domains such as gene regulatory analysis and biomedical research. Existing causal discovery methods either learn a global causal structure, which incurs substantial computational cost, or assume the absence of latent variables and selection bias, assumptions that are often violated in real-world settings. Motivated by these challenges, we study local causal structure learning in the presence of latent variables and selection bias. Specifically, we first characterize a local region that enables target-specific causal discovery without recovering the entire global structure. We then establish a theoretical bridge between causal information learned from the observed distribution induced on this local region and the corresponding information in the global causal structure. Building on these foundations, we propose LoCaLS, a local causal structure learning algorithm that is sound and complete under standard assumptions and identifies the same direct causes and effects of a target variable as those identifiable by global causal discovery methods, while allowing for latent variables and selection bias. Extensive experiments on random and real-world structures demonstrate that the proposed method consistently achieves higher structural accuracy than existing local methods while requiring substantially less computational effort than state-of-the-art global methods. Furthermore, applications to two real-world gene expression datasets reveal biologically plausible target-specific causal structures, demonstrating its practical applicability in large-scale biological data analysis.

View source

Similar papers

Preprint Aug 2026

Interpretable Causal Discovery via Causal-Effect Constraints

This work considers the task of conditional causal discovery as a Bayesian inference problem, in which the posterior is targeted over causal graphs and parameters conditional on an event such as a causal-effect constraint, and adapts rare-event estimation techniques to perform inference the joint graph-parameter space.

Cixuan Zhang, Guy Van den Broeck, Benjie Wang · 0 citations
Book Open access Aug 2026

Continuous Causal Component and Structure Discovery from Time Series

Causal representation learning aims to infer a small set of causally related latent variables and to model their causal relationships in order to explain high-dimensional observed data. Recent studies have made progress in identifying causal representations from time series by hypothesizing temporal structures among latent components. However, these methods are constrained by the observation scale and sampling regularity, as they typically assume that latent components evolve at discrete time steps. Moreover, they do not explicitly identify the causal graph among latent components, which limits the interpretability of the learned representations. To address these limitations, we propose Continuous Causal Component and Structure Discovery (C3SD). We theoretically show that latent causal components and their causal relationships can be identified up to permutation equivalence by modeling synchronous sparsity in the mapping between latent components and observed variables. Building on this result, C3SD employs a dual sparsity-induced autoencoder to infer latent causal components, together with an adaptive group lasso to jointly structure the encoding and decoding matrices. In addition, a neural ordinary differential equation–based joint autoencoder models the continuous-time causal dynamics of the latent components and recovers their underlying dynamical causal structure. Extensive experiments demonstrate that C3SD effectively identifies latent components and the causal mechanisms driving their continuous temporal evolution, particularly in sparsely and irregularly sampled time series.

Dezhi Yang, Jun Wang, C. Domeniconi et al. · 0 citations
Preprint Aug 2026

Joint Causal Structure and Cluster Discovery Using Variational Inference

A novel approach based on variational inference to simultaneously infer both the latent clusters and causal structures is presented and an approximate posterior over clusters and graph-structure is learned by considering variational distributions based on categorical and Bernoulli models.

Avni Rajpal, Anubhav Kumar, Rishabh Karnad et al. · 0 citations
Open access Jul 2026

Local Causal Structure Learning with Efficient Parents Discovery

Local causal structure learning aims to identify direct causes (parents) and effects (children) of a target variable. Recent advances in this field rely on learning the MB (Markov Blanket) of unidentified variables. However, most existing methods require extensive search to discover spouses during MB learning, and they often need to learn the MB within the PC (parents and children) set of a target variable to identify parents, which becomes computationally expensive when the parent set is large and complex. To address this issue, we propose a novel local causal structure learning algorithm with Efficient Parents Discovery, named EPD. Specifically, EPD introduces an MB discovery subroutine, MBDis, which first identifies some parents of the target variable using the V-structure, and then performs feature selection to exclude candidate spouses that are weakly related to the target variable, thereby reducing the size of the candidate spouse set and accelerating spouse discovery. Additionally, EPD incorporates an IdePC subroutine, which learns the PC sets of the unresolved variables to identify additional parents, reducing the search space of parents. With the proposed MBDis and IdePC subroutines, EPD adopts an MB-by-PC learning strategy; it starts from discovering the MB of the target variable and then learns the PC sets of undetermined variables. This process continues iteratively until the parents and children of the target variable are identified. Using 7 Bayesian networks and 1 real-world data, the experiments have verified the effectiveness of EPD, in comparison with 10 state-of-the-art methods.

Shuai Yang, Xin-Yu Miao, Xianjie Guo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Symmetries and Causality: Causal Effect Identification Beyond IID Data

In the natural sciences, symmetries and cause-effect relationships are ubiquitous. Yet for complex machine-learning tasks, like world-modeling in reinforcement learning, they appear difficult to harness. We propose a formal description of statistical systems based on symmetries in data leaving causal mechanisms invariant. The result is an abstract, simple and general mathematical language for causal reasoning. This paper provides formal descriptions of models and queries, setting up this language, and the formal infrastructure and strategies for their mathematically rigorous identification from data within this formalism. This approach reproduces and matches standard theoretical results on IID data and transport of experimental and non-experimental data. But its main purpose is to unify and substantially extend the scope of causal reasoning, in going beyond IID data and in approaching complex causal queries not captured by do- or soft-interventions. This new perspective on causally relevant aspects of data-modeling additionally sheds new light on well-known structures like c-components or hedges but also includes aspects of missing data and is inherently well-suited for the description of transfer and robustness properties.

Martin Rabel, Jakob Runge · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.