Missing data constitute a pervasive challenge in empirical research. Consequently, there is an ever-growing number of methods designed to address this challenge, with multiple imputation and inverse probability weighting the dominant strategies. Despite this, theoretical guarantees remain limited, particularly in the challenging case of non-monotone missing at random (MAR). When guarantees exist, they are often confined to simplified settings such as missing completely at random, monotone or block-wise missingness, or rest on restrictive assumptions about the missingness mechanism. In this paper, we utilize the theory of sieve maximum likelihood to establish a general rate of convergence under MAR that requires no modeling of the missingness mechanism and no restriction on the configuration of missing patterns, beyond MAR itself and a natural positivity condition. Applying this result to density estimation, we show that the complete-data density can be estimated at the minimax rate over a H\"older class, up to a logarithmic factor, for any prescribed smoothness level. The missingness does not affect the rate and enters only through a constant. The estimator is approximated in practice by a simple expectation-maximization (EM) algorithm operating on the incomplete data directly. In simulations, it performs comparably to the kernel density estimator supplied with the complete data across a wide range of missingness levels.
Digital twins for health may be used to compare treatments, project patient trajectories, and support clinical decisions. While related to mechanical digital twins, those initially developed for engineering applications, replicating the mechanical digital twin architecture and goals may fail in health for two reasons. The fidelity trap is the belief that an accurate model can answer what-if questions by virtue of its accuracy. Prediction and counterfactual reasoning are different tasks, and a twin that can fit past trajectories well may miss the mark when ranking treatments. The feedback trap arises when the twin updates on data its own recommendations helped generate. Refitting in this way can recover a biased relationship and grow more confident even as data grows thinner. We contend that health digital twins should be conceived as causally valid, modular, and evolving systems. Modularity isolates the data and models needed for interventional recommendations, causal validity supports such claims, and governed evolution updates the twin while accounting for how its recommendations reshape the data. We conclude that the standard for a health twin should be how well it supports decisions in the world it helps create, not how faithfully it reproduces the world it observes.
Nikki L. B. Freeman, Yating Zou, Kyungbok Lee et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.