Individualized treatment rules (ITRs) map baseline characteristics to treatment recommendations, with the optimal ITR maximizing expected reward or policy welfare. Indirect methods may require restrictive modeling assumptions, whereas direct methods can be sensitive to nuisance estimation error and limited overlap. We propose orthogonal double residual learning (ODRL), a two-stage, cross-fitted framework that directly targets the optimal ITR through cost-sensitive classification using the product of treatment and outcome residuals. To our knowledge, ODRL is the first direct method with a universally Neyman orthogonal objective requiring neither restrictive modeling assumptions nor inverse propensity score weighting. Thus, nuisance estimation errors affect regret through a second-order product, and ODRL remains robust under limited overlap. The Fisher consistent objective accommodates general decision rule sieves. We establish nonasymptotic high probability value function regret bounds relative to the Bayes classifier for VC classes, including linear rules and decision trees, and calibrated regret bounds for surrogate relaxations using support vector machines and deep ReLU neural networks. We further show that generic surrogate relaxations need not preserve orthogonality, whereas bounded score hinge learning does. Simulations demonstrate strong performance across complex and linear decision boundaries, limited overlap, and working model misspecification. Applications to the Right Heart Catheterization study and the Oxford Net Zero experiment illustrate interpretable treatment or policy recommendations. The \texttt{odrlITR} R package implements ODRL.
We develop a partial identification learning framework for individualized treatment rules (ITRs) with categorical treatments, outcomes, and instrumental variables. Rather than relying on strong causal assumptions required for point identification, our framework leverages causal bounds to characterize the optimal treatment decision. Existing methods for ITR optimization under partial identification are largely restricted to binary treatment settings and the bounds derived by Balke and Pearl under the canonical instrumental variable design. We extend this framework to accommodate a broader class of causal structures as well as scenarios with categorical treatment, outcome, and instrumental variables. We introduce a generalized minimax loss criterion for treatment selection from among more than two options, which minimizes the maximum possible difference between the chosen and the optimal treatment based on partial identification bounds. To construct the ITR, we use a symmetric embedding strategy that maps discrete treatments to the vertices of a regular simplex, avoiding the geometric inconsistencies of standard one-vs-rest approaches. We derive a differentiable, weighted surrogate risk function and show that optimizing it solves the original problem. Furthermore, we provide finite sample convergence rates via an oracle inequality under general regularity conditions, which we show are satisfied by a kernel based implementation. Numerical experiments demonstrate that the framework yields ITRs significantly closer to the oracle ITR compared to existing alternatives in settings with unmeasured confounding.
Johannes Hruza, Paweł Morzywołek, Jakob Zeitler et al.· 0 citations
There is an increasing call for individualized treatment rules, which leverage individual patient characteristics to recommend treatments or interventions, tailoring recommendations based on their covariates. This is particularly of interest for the care of conditions such as depression, for which many treatment options are available with similar average effectiveness but with large heterogeneity in individual responses. In parallel, there has been a growing interest in machine learning methods for causal inference and variable selection. We compared several strategies for variable adjustment in a dynamic marginal structural modeling approach to estimating an optimal individualized treatment rule and investigated the performance of Outcome Adaptive Lasso, Group Lasso and Doubly Robust Estimation, Double-index Propensity Score, the High-Dimensional Balancing Propensity Score, and the Causal Ball Lasso as variable selection methods for the propensity score. Our results demonstrate that these all provided similar unbiased estimates. However, methods differed in their ability to exclude extraneous variables and in computational burden. We found statistical efficiency is gained when variable selection approaches were for the propensity score were used and by including variables in the outcome model. We applied all methods to determine the optimal treatment rule, treating with either selective serotonin reuptake inhibitors or serotonin and norepinephrine reuptake inhibitors, for unipolar depression in individuals aged 13 years and older. This analysis, which used electronic health records from 74,058 Kaiser Permanente Washington patients with a new antidepressant dispensing between 2008 and 2018, suggested tailoring treatment based on baseline symptom severity did not impact symptom severity 6 months later.
N. Galanter, S. Shortreed, Erica E. M. Moodie· Psychological methods· 0 citations
In marketing, optimizing subsidy allocation to maximize overall profits is of substantial economic importance. Prior research has employed treatment effect estimation techniques to identify subsidy-sensitive items and design corresponding allocation strategies. However, more accurate treatment effect estimations do not necessarily lead to better allocations, underscoring the critical influence of decision boundaries in decision-making. This paper argues that optimal allocation fundamentally depends on predicting the expected optimal subsidy, a challenge distinct from conventional treatment effect estimation or causal decision-making, which existing approaches fail to address. To fill this gap, we introduce a two-stage Counterfactual optimal subsidy Learning method with an Asymmetric reward (CoLA). In the first stage, we derive a coarse estimate of the expected subsidy threshold by exploiting order information and the conditional independence between expected and observed subsidies. In the second stage, we refine these estimates using an asymmetric loss function, leading to more robust predictions. Under practical budget constraints, we prioritize candidates based on their Sharpe ratios to determine the final subsidy allocation strategy. Experiments on three public datasets and an online A/B test show that our method achieves significant performance improvements, yielding the highest total profit and incremental leverage ratios.
Xiang Li, Yanghao Xiao, Chunyuan Zheng et al.· Annual International ACM SIG...· 2 citations
Structured pruning uses surrogate objectives because direct task evaluation over every feasible mask is too expensive. Most evaluations report average surrogate error or rank correlation on broadly sampled masks. These summaries do not directly test the mask chosen by the surrogate. We introduce PruneShift, an evaluation framework that separates broad predictive fidelity, fidelity near selector outputs, and the quality of the selected pruning decision. We first prove that Spearman and Kendall agreement can approach one while normalized selection regret remains maximal. We then derive sufficient conditions based on uniform error, selector suboptimality, decision margin, density ratio, and comparison mass. The analysis also yields a finite pool certificate with an explicit excess cost bound. Four studies test different links in this argument. External TextbookQA confirmation is heterogeneous: 7 of 20 simultaneous intervals favor the surrogate-selected mask, 6 favor its fixed comparator, and 7 cross zero. On a fixed Natural Questions pool, strict improvement holds in one of four settings. A controlled QQP experiment supports the proposed coverage mechanism in all 16 prespecified endpoints, although the sufficient bounds are conservative. Finally, a restricted OSSCAR reconstruction study on OPT-125M shows better local than broad fidelity in 68 of 75 primary endpoints. Independent fixed-mask confirmation is inconclusive in 24 of 25 endpoints and favors the comparator in one. These results show why predictive fit, decision reliability, and pruning method quality require separate evidence.
This work derives a finite-dimensional dual formulation of PrO inference that separates sampling fluctuation, approximation under a divergence budget, regularization, and numerical optimization error and uses an exactly solvable categorical example to show that predictive-risk convergence can imply convergence to a unique predictive distribution even though the parameter distributions have no weak limit on the original parameter space.
Aurya Javeed, D. Kouri, Teresa Portone et al.· 0 citations
Finding the optimal individualized treatment rule that maps individual characteristics or contextual information to treatment assignments has been extensively investigated in existing literature, with widespread practical applications. This paper considers the estimation of optimal treatment regimes within a semi-supervised data framework (exemplified by electronic medical record data). In such settings, only a tiny proportion of observations have observed outcome labels, owing to high labeling costs, time limitations, data privacy concerns, and other constraints, while covariates and treatment assignments are available for all study subjects. We develop a semi-parametric inference method for optimal treatment regimes, which leverages outcome- unlabeled samples with complete covariate and treatment information to enhance estimation efficiency. The proposed estimation framework consists of two key steps: first, flexible nonparametric imputation via single-index kernel smoothing; second, subsequent estimation of the optimal treatment regime based on concordance-assisted learning. We establish the consistency and asymptotic normality of our proposed estimators. Numerical simulation studies demonstrate that our method achieves higher efficiency and stronger robustness relative to fully supervised estimators under finite-sample settings. We further validate the practical value of our proposed framework using the MIMIC-III and ACTG175 datasets.
Mengjiao Peng, Yong Zhou, Wenbin Lu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.