Counterfactual Optimization of Policy Interventions: Lexical Ordering and Leapfrogging
Most data-driven policy learning methods maximize average outcomes, overlooking the possibility that a policy beneficial on average may still harm a substantial fraction of individuals. Motivated by the ethical principle of"first do no harm", we study how to design a change from a baseline policy that improves overall...