Offline Policy Evaluation and Learning with Harm Constraints
Offline policy learning aims to optimize individualized decisions using historical data. However, conventional methods primarily focus on maximizing expected rewards while neglecting individual-level counterfactual harm—cases where the assigned treatment leads to worse outcomes than the control. This can result in over...