Most data-driven policy learning methods maximize average outcomes, overlooking the possibility that a policy beneficial on average may still harm a substantial fraction of individuals. Motivated by the ethical principle of"first do no harm", we study how to design a change from a baseline policy that improves overall welfare while keeping the worst-case probability or expectation of individual harm below a specified limit. We establish sufficient conditions under which an optimal policy transition has a lexical leapfrogging structure: groups defined by covariates and current treatment are ranked by a priority score, and any treatment change moves them directly to the conditionally optimal treatment. We derive this score under several models for the dependence among potential outcomes. We demonstrate this harm-aware policy optimization approach in a reanalysis of the I-SPY2 breast cancer platform trial and show how the consideration of counterfactual harm may lead to different conclusions about which treatment-subgroup pairs may warrant deprioritization in further clinical evaluation.
Offline policy learning aims to optimize individualized decisions using historical data. However, conventional methods primarily focus on maximizing expected rewards while neglecting individual-level counterfactual harm—cases where the assigned treatment leads to worse outcomes than the control. This can result in over...
Jile Chaoge, Qin-Wei Yang, Jing-Yi Li et al.· Entropy· 0 citations
Policy learning methods based on conditional average treatment effects can obscure subpopulation heterogeneity when applied to ordinal outcomes. We develop a policy learning framework for ordinal outcomes with heterogeneous utilities for individuals who strictly benefit from treatment and those who do not. Since the pr...
This paper studies risk-averse treatment allocation when individuals self-select into treatment based on unobserved characteristics. We develop a framework that combines the marginal treatment effect approach to endogenous selection with a general class of coherent risk measures that capture distributional preferences...
Reinforcement learning (RL) seeks to optimize sequential decisions to maximize population-level benefits over time. However, when deployed in high-stakes settings such as healthcare, RL decisions might systematically restrict some subpopulation's access to valuable services in a manner contrary to the values and goals...
Jianhan Zhang, Jitao Wang, John D. Piette et al.· 0 citations
A central goal when designing treatment policies is often to"do no harm", that is, to avoid interventions that improve average outcomes while worsening outcomes for some individuals. A widely used notion for harm is the fraction of negatively affected (FNA), defined as the probability that an intervention decreases an...
Rui-Zi Yan, Dennis Frauen, Maresa Schröder et al.· 0 citations
Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created by the intervention itself. We develop this theory for learning systems, where states equ...
Qin-You Wang· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.