A Data-Centric Decomposition of Estimator Performance in Continuous Treatment Effect Estimation
Abstract
Standard supervised learning algorithms prioritize predictive performance over causal inference, making them ill-suited for conditional average dose response (CADR) estimation. Although specialized CADR estimators have been proposed to address this, the field lacks clarity on which dataset characteristics truly challenge model performance, largely because current benchmarking practices fail to isolate different sources of estimation error. In this work, we analyze these current benchmarking practices and introduce a novel decomposition framework that disentangles the contribution of distinct data-generating components, such as confounding, dose distribution non-uniformity, and response surface complexity, to estimator performance. Applying this scheme to several established benchmarks, we uncover that widely used datasets do not primarily test for confounding robustness, as often assumed, but are instead dominated by challenges arising from non-uniform dose distributions. We further propose a new benchmark dataset with high CADR heterogeneity, where confounding does have a substantial impact. Our results call for a rethinking of current evaluation practices and advocate for more diagnostic, data-centric benchmarks to advance the development of robust CADR estimation methods.