Policy learning methods are increasingly used to inform treatment allocation under budget constraints. Most proposed methods assume complete treatment data, yet applications frequently suffer from missingness that can bias estimates and lead to suboptimal policies. We address this gap by extending efficient estimators for average treatment effect (ATE) estimation to policy value and conditional average treatment effect (CATE) estimation under missing at random (MAR) and missing completely conditionally at random (MCCAR) treatment data. Through asymptotic efficiency analysis, we prove that the MAR estimator, which leverages partially-observed units, is both valid and more efficient than the MCCAR estimator when MCCAR assumptions hold. This result provides formal justification for preferring MAR-based estimation in policy learning under both missing data settings. Our comprehensive experiments using synthetic and semi-synthetic datasets confirm that correctly specifying the missingness mechanism is crucial: misspecified estimators remain biased regardless of sample size, while our estimators achieve near-oracle performance when assumptions are satisfied. Our work provides practitioners with theoretically grounded, empirically validated tools for robust policy learning in the presence of missing treatment data.
Estimating the average treatment effect (ATE) remains a fundamental challenge in observational studies in the presence of poor or limited covariate overlap. Although the inverse probability weighting (IPW) estimator is a widely used approach for estimating the ATE, its performance can deteriorate substantially when overlap is limited, often resulting in increased finite sample bias and unreliable confidence intervals. One common strategy is to shift attention from the original target estimand, the ATE, to alternative estimands that are less sensitive to extreme propensity scores; however, doing so changes the scientific question of interest. In this manuscript, we propose a novel ATE estimator that preserves the original target estimand, the ATE, while improving robustness to limited overlap. A key idea is that a class of estimands can be expressed by a polynomial function of a hyperparameter characterizing the estimands. Exploiting this structure, the proposed method computes IPW estimators for a sequence of such estimands, models these estimates using a polynomial function, and extrapolates to recover the ATE. We show that the estimator has consistency and asymptotic normality under weaker overlap conditions than required for the standard IPW estimator. Simulation studies demonstrate that the proposed method improves estimation accuracy and interval performance in settings with limited overlap. In addition to its theoretical and empirical advantages, the proposed approach has a clear interpretation and is easy to implement using standard statistical software.
Shunichiro Orihara, S. Komukai, Fan Li· 0 citations
Randomized controlled trials (RCTs) are fundamental tools for causal inference across technology companies, pharmaceutical research, and federal agencies. While the standard difference-in-means estimator provides unbiased treatment effect estimates, it often lacks precision, particularly when treatment effects are heterogeneous or outcomes exhibit heavy-tailed distributions. Although numerous precision-enhancing methods exist---from covariate adjustment techniques to variance reduction strategies---recent research demonstrates that no single estimator performs optimally across all datasets. Rather than seeking the best estimator for individual RCTs, which risks compromising scientific validity through convenient selection, we propose a principled framework for identifying optimal estimators within families of RCTs based on specific analytical goals. Our approach uses sample splitting to estimate the distribution of evaluation metrics (e.g., mean squared error, regret) across RCT families, enabling systematic comparisons between estimators while maintaining asymptotic guarantees. We demonstrate this framework using a sample of Amazon's Supply Chain Optimization Technology trials and the Strengthening Democracy Challenge dataset (25 interventions). Results reveal that optimal estimators vary significantly by analytical objective: weighted least squares performs best for inference goals, while difference-in-means minimizes regret for decision-making contexts. This work provides actionable guidance for estimator selection while preserving methodological rigor across diverse research applications.
Harsh Parikh, Gabriel Levin-Konigsberg, Nilesh Tripuraneni et al.· 0 citations
Individualized treatment rules (ITRs) map baseline characteristics to treatment recommendations, with the optimal ITR maximizing expected reward or policy welfare. Indirect methods may require restrictive modeling assumptions, whereas direct methods can be sensitive to nuisance estimation error and limited overlap. We propose orthogonal double residual learning (ODRL), a two-stage, cross-fitted framework that directly targets the optimal ITR through cost-sensitive classification using the product of treatment and outcome residuals. To our knowledge, ODRL is the first direct method with a universally Neyman orthogonal objective requiring neither restrictive modeling assumptions nor inverse propensity score weighting. Thus, nuisance estimation errors affect regret through a second-order product, and ODRL remains robust under limited overlap. The Fisher consistent objective accommodates general decision rule sieves. We establish nonasymptotic high probability value function regret bounds relative to the Bayes classifier for VC classes, including linear rules and decision trees, and calibrated regret bounds for surrogate relaxations using support vector machines and deep ReLU neural networks. We further show that generic surrogate relaxations need not preserve orthogonality, whereas bounded score hinge learning does. Simulations demonstrate strong performance across complex and linear decision boundaries, limited overlap, and working model misspecification. Applications to the Right Heart Catheterization study and the Oxford Net Zero experiment illustrate interpretable treatment or policy recommendations. The \texttt{odrlITR} R package implements ODRL.
I study the optimal design and analysis of randomized experiments for estimating finite-population average treatment effects when potential outcomes are known to be bounded, as with binary outcomes. Among all assignment mechanisms and a broad class of affine estimators, worst-case mean-squared error (MSE) is minimized by independent random assignment and an unconventional regression of the support-midpoint-centered outcome on the recentered treatment, with no intercept. This contrasts with the usual prescription of balanced complete randomization and difference-in-means estimation: when outcomes are bounded, randomness in the realized treatment share is informative. The worst-case gain over full-sample complete randomization is asymptotically small, but gains can be first-order relative to other designs: complete within-pair randomization and pair-fixed-effect regression have twice the worst-case MSE. I extend the result to allow for arbitrary estimators. Independent random assignment remains optimal, and the generally-nonlinear optimal estimator can meaningfully reduce worst-case MSE.
Much causal inference research is focused on methods for optimizing dynamic treatment regimes (Murphy, 2003; Robins, 2004; Schulte et al., 2015), which are rules for deciding which treatments should be assigned and when based on evolving history. There is a certain optimism underlying this endeavor that with enough tinkering we might realize consequential improvements. Another strand of research, previously confined to the point exposure setting, considers bounds on how well any individualized treatment rule could possibly do. Here, we extend to the time-varying setting sharp bounds on the performance of an oracle strategy that selects the best treatment regime for each subject based on their unobserved potential outcomes or `response type'. For binary outcomes, the lower bound (assuming higher is better) is simply the expected outcome attained by the optimal treatment regime based on observed history. For continuous outcomes, the lower bound may strictly exceed the maximal observed covariate based value. In the continuous setting, we also consider bounds on the CDF of oracle continuous potential outcomes.
Self-supervised Causal Effects Estimation is proposed, a novel framework that integrates causal priors with self-supervised learning to construct balanced and predictive representations for causal effects estimation that consistently outperforms state-of-the-art methods.
Xin-Shu Li, Shiyi Yang, Venus Haghighi et al.· ACM Transactions on Intellig...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.