Skip to content
Open access

A Copula-Tensor Neural Network Framework for High-Dimensional Causal Inference

Aug 2026 · Mathematics · 0 citations · 13 references

TL;DR

It is suggested that copula-based feature representations combined with deep learning provide a flexible approach for heterogeneous treatment effect estimation, particularly in high-dimensional settings with complex covariate dependence.

Abstract

Estimating conditional average treatment effects (CATEs) in high-dimensional causal inference problems remains challenging because complex nonlinear relationships, heterogeneous feature distributions, and dependence among covariates can limit the effectiveness of conventional machine learning approaches. To address this challenge, we propose a copula-enhanced neural learning framework that integrates empirical copula transformations, manifold-based feature augmentation, structured treatment–covariate interaction representations, and deep neural networks for flexible CATE estimation. The empirical copula transformation does not introduce additional dependence information; instead, it provides a rank-based feature representation that normalizes marginal distributions, reduces sensitivity to heterogeneous feature scales and extreme observations, and offers a dependence-aware representation for subsequent learning. The proposed framework is evaluated through Monte Carlo simulations under diverse data-generating mechanisms and a real-world application using the Criteo uplift dataset. The simulation study examines the contribution of individual model components through ablation experiments and compares the proposed approach with established causal learning methods. Results demonstrate that the proposed framework achieves competitive CATE estimation accuracy while providing stable policy evaluation based on Inverse Propensity Scoring (IPS) and Doubly Robust (DR) estimators. In the Criteo application, the proposed method exhibits predictive performance comparable to conventional neural-network approaches while producing more stable Doubly Robust policy value estimates. These findings suggest that copula-based feature representations combined with deep learning provide a flexible approach for heterogeneous treatment effect estimation, particularly in high-dimensional settings with complex covariate dependence. The benefits of the proposed framework depend on data characteristics, including sample size, dimension, dependence structure, and treatment assignment mechanisms.

Read PDF

Similar papers

#machine learning Preprint Aug 2026

A Deep Latent Variable Framework for Jointly Modeling Missingness, Measurement Error, and Heterogeneity

A unified probabilistic framework that jointly addresses missing data, measurement error, and population heterogeneity utilizing deep latent variable representation is proposed that integrates a novel hierarchical tree-routed variational autoencoder with pattern-aware latent representations and calibration-based denoising.

Yasin Khadem Charvadeh, Grace Y. Yi, Mithat Gönen et al. · 0 citations
#machine learning Preprint Sep 2026

Representation Learning for Sample-Efficient CATE Estimation by Leveraging Multiple Outcomes

Estimating conditional average treatment effects (CATE) enables efficient targeting of interventions, but many applications have limited experimental samples, making it difficult to estimate heterogeneous effects from high-dimensional covariates. In such settings, policymakers and medical practitioners often succumb to the curse of dimensionality or apply off-the-shelf dimension reduction methods that may not preserve treatment heterogeneity. Yet these domains often come with large historical datasets measuring a wide range of outcomes -- a source of supervision that is rarely exploited in practice. Following causal representation learning, we hypothesize that such domains with high-dimensional covariates have lower-dimensional underlying dynamics. We can thus leverage the diverse outcomes measured in historical data to learn a lower-dimensional representation of the covariates. Theoretically, we prove that when the auxiliary outcomes satisfy a set of surrogacy conditions and the representation retains relevant covariate information, the original CATE is identified when the high-dimensional covariates are replaced by the learned representation. Combined with existing dimension-dependent rates for CATE estimation, the result implies greater sample-efficiency on the same experimental sample. Additionally, we characterize the bias-variance tradeoff when the assumptions do not hold perfectly, and show that the representation-based estimator can still achieve lower error when the reduction in estimator variance outweighs the bias due to compression. Empirically, we evaluate the method on synthetic data and semi-synthetic medical data.

Maitreyi Swaroop, Shikha Bhat, Samantha Rodriguez et al. · 0 citations
Preprint Sep 2026

From Good Starts to Optimal Inference: Generalized Latent Factor Models with Missingness and Implicit Regularization

Generalized latent factor models provide a flexible framework for analyzing high-dimensional non-Gaussian data, but principled estimation and uncertainty quantification under missingness remain substantially less developed. We develop a theory that connects a computationally tractable nonconvex procedure directly to statistical inference for nonlinear latent factor models with exponential-family links and partially observed entries. Our procedure combines a link-aware double-SVD initialization, a unilateral refinement that achieves rowwise consistency, and vanilla gradient descent. We show that the refined initializer enters a region of incoherence and contraction and that gradient descent remains in this region through implicit regularization, contracting rapidly down to the statistical estimation error without explicit incoherence or balancing regularization. Our central result is a uniform rowwise linear approximation for the actual output of gradient descent that isolates the leading score fluctuations from higher-order estimation and optimization errors. These expansions yield asymptotically valid individual and Gaussian multiplier-bootstrap simultaneous inference for latent factors, together with simultaneous confidence bands for missing-entry means, without requiring an additional debiasing step. The resulting estimation rate matches a restricted-class minimax lower bound up to logarithmic factors, while the theory accommodates severe missingness, weak low-rank signals, and diminishing local curvature. Simulations support the theoretical findings, and an application to large language model evaluation illustrates uncertainty-aware estimation and ranking of latent model capabilities.

Cheng-Zhu Huang, Yu-Qi Gu · 0 citations
Jul 2026

Variational meta-learning inference for low dimensional neural system identification

This work proposes a fully probabilistic extension of the manifold meta-learning framework, based on amortized Variational Inference, where a generative prior over the low-dimensional parameter manifold is learned.

Matteo Rufolo, D. Piga, M. Forgione · 0 citations
Book Open access Aug 2026

Sparse Additive Models for Domain Generalization

Machine learning models continue to face challenges in out-of-distribution (OOD) generalization, where domain generalization (DG) aims to improve performance on unseen domains under distributional shifts. A prevalent paradigm in DG focuses on learning domain-invariant feature representations. However, feature representations from existing methods often exhibit weak interpretability. To bridge this gap, we propose Sparse Additive Domain Generalization (SpADG). We incorporate an additive structure into the DG framework and employ ℓq,1 -norm regularization to induce sparsity, thereby enabling structured feature selection and enhancing interpretability. We present two distinct realizations: an additive kernel-based formulation and a neural additive model-based approach. The former leverages the representer theorem for flexible data adaptation, while the latter learns nonlinear shape functions. Theoretically, we derive generalization error bounds for both realizations and prove the feature selection consistency of our method under rate-scaled regularization condition. Empirical evaluations on synthetic and real-world datasets validate the effectiveness of SpADG, particularly its robustness in high-dimensional settings.

Jiayi Wang, Han Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.