Skip to content

Distributionally Faithful Imputation via Positive Semi-Definite Kernel Density Estimation

Jul 2026 · arXiv.org · Vol abs/2607.07767 · 0 citations · 37 references
Mathematics Computer Science

TL;DR

This work recast imputation under missing completely at random (MCAR) as density estimation from masked observations: estimate a distribution whose observed marginals exactly match those in the data.

Abstract

Missing values undermine statistical inference and machine learning pipelines, yet most imputation methods rely on heuristics or restrictive parametric assumptions that ignore the joint data distribution. We recast imputation under missing completely at random (MCAR) as density estimation from masked observations: estimate a distribution whose observed marginals exactly match those in the data. Leveraging positive semi definite (PSD) kernel densities we obtain a convex empirical risk problem with closed form marginals, solvable by a Newton interior point method. The resulting PSD Impute model yields both single and multiple imputations from the same fitted density, enjoys statistical consistency with fast adaptive excess risk beating the curse of dimensionality for very regular probabilities. Preliminary experiments on one synthetic and eleven real world datasets already indicate competitive distributional accuracy compared with popular imputation baselines, suggesting strong practical promise.

View source

Similar papers

Preprint Aug 2026

Asymptotics of Nonparametric Estimation under General Non-monotone MAR Missingness: A Nonparametric Maximum Likelihood Approach

Missing data constitute a pervasive challenge in empirical research. Consequently, there is an ever-growing number of methods designed to address this challenge, with multiple imputation and inverse probability weighting the dominant strategies. Despite this, theoretical guarantees remain limited, particularly in the challenging case of non-monotone missing at random (MAR). When guarantees exist, they are often confined to simplified settings such as missing completely at random, monotone or block-wise missingness, or rest on restrictive assumptions about the missingness mechanism. In this paper, we utilize the theory of sieve maximum likelihood to establish a general rate of convergence under MAR that requires no modeling of the missingness mechanism and no restriction on the configuration of missing patterns, beyond MAR itself and a natural positivity condition. Applying this result to density estimation, we show that the complete-data density can be estimated at the minimax rate over a H\"older class, up to a logarithmic factor, for any prescribed smoothness level. The missingness does not affect the rate and enters only through a constant. The estimator is approximated in practice by a simple expectation-maximization (EM) algorithm operating on the incomplete data directly. In simulations, it performs comparably to the kernel density estimator supplied with the complete data across a wide range of missingness levels.

Yating Zou, Huimin Hu, Jeffrey Näf · 0 citations
Preprint Jul 2026

Handling Missingness and Censoring in Dirichlet Mixture Models

Incomplete compositional data analysis faces a fundamental limitation: likelihood-based methods for compositional models generally require fully observed compositions, making it difficult to accommodate missing or censored proportions directly on the simplex. Consequently, analysts often discard partially observed compositions or transform the data into unconstrained spaces, potentially sacrificing interpretability and coherence. This paper proposes a likelihood-based method for incomplete compositional data without leaving the simplex. Specifically, we develop an Expectation-Maximisation (EM) type algorithm for fitting finite mixtures of Dirichlet distributions in the presence of missing and censored components. The proposed approach performs parameter estimation and model-based imputation simultaneously while preserving the compositional structure and interpretability of the original variables. A simulation experiment evaluates the performance of the proposed estimators and imputations under increasingly complex coarsening mechanisms. Particular attention is paid to clustering performance, and model selection outcomes. The results showed beneficial clustering performance despite observations being incomplete, and a higher probability of model selection metrics identifying the correct number of clusters compared to current alternative of case-deletion. The practical utility of the method is illustrated using two real datasets with distinct coarsened patterns. Analysis of the xenolith dataset identifies a four-component Dirichlet mixture that reveals interpretable profiles of rock types and speciation methods. Application to PM$_{2.5}$ speciation data from the Air Quality System, containing both left-censored and missing-at-random values, supports a four-component mixture model that characterises compositional parts of particulate matter across the United States.

J. Pillay, A. Bekker, C. Tortora et al. · 0 citations
Preprint Sep 2026

From Good Starts to Optimal Inference: Generalized Latent Factor Models with Missingness and Implicit Regularization

Generalized latent factor models provide a flexible framework for analyzing high-dimensional non-Gaussian data, but principled estimation and uncertainty quantification under missingness remain substantially less developed. We develop a theory that connects a computationally tractable nonconvex procedure directly to statistical inference for nonlinear latent factor models with exponential-family links and partially observed entries. Our procedure combines a link-aware double-SVD initialization, a unilateral refinement that achieves rowwise consistency, and vanilla gradient descent. We show that the refined initializer enters a region of incoherence and contraction and that gradient descent remains in this region through implicit regularization, contracting rapidly down to the statistical estimation error without explicit incoherence or balancing regularization. Our central result is a uniform rowwise linear approximation for the actual output of gradient descent that isolates the leading score fluctuations from higher-order estimation and optimization errors. These expansions yield asymptotically valid individual and Gaussian multiplier-bootstrap simultaneous inference for latent factors, together with simultaneous confidence bands for missing-entry means, without requiring an additional debiasing step. The resulting estimation rate matches a restricted-class minimax lower bound up to logarithmic factors, while the theory accommodates severe missingness, weak low-rank signals, and diminishing local curvature. Simulations support the theoretical findings, and an application to large language model evaluation illustrates uncertainty-aware estimation and ranking of latent model capabilities.

Cheng-Zhu Huang, Yu-Qi Gu · 0 citations
Preprint Jul 2026

From dense grids to valid inference: Accounting for regularization bias in nonparametric random coefficient models

This paper develops an inference procedure for average functionals of random-coefficient distributions, such as mean willingness-to-pay and average elasticities, when the distribution is estimated nonparametrically using the penalized fixed-grid estimator of Heiss, Hetzenecker, and Osterhaus (2022). We establish asymptotic normality of the corresponding penalized plug-in estimator centered at the functional evaluated at the penalized pseudo-true value and propose a confidence interval that accounts for the regularization bias. Our method applies to a broad class of linear and nonlinear functionals and allows researchers to use dense grids to reduce approximation bias while maintaining valid inference. Monte Carlo simulations show that the proposed intervals achieve coverage close to the nominal level while remaining informative in finite samples. An empirical application to travel mode demand illustrates that flexible nonparametric specifications can yield economically meaningful differences relative to standard parametric models.

Ling-Yan Kong, M. Osterhaus, Michael Pen · 0 citations
Preprint Aug 2026

Handling Missing Data in Probabilistic Regression Trees

Probabilistic Regression Trees (PRTrees) are a smooth and consistent alternative to classical regression trees, producing continuous predictions through probabilistic split assignments. This paper extends the PRTree framework to accommodate missing predictor values directly during tree construction, eliminating the need for prior imputation. Three strategies are proposed, each exploiting the available information differently: a uniform-probability approach, a partial-observation approach, and a dimension-reduced smoothing approach. These modifications are defined to preserve the fundamental probabilistic properties of the original methodology, including probability conservation and marginal compatibility, under arbitrary patterns of missing covariate values. The proposed methods are evaluated on several real-world datasets exhibiting different levels of missingness and are compared with classical regression trees. The results show that the effectiveness of probabilistic tree construction depends strongly on the treatment of missing observations. Across the considered datasets, the fill strategy emerged as the dominant modeling component, often exerting a larger influence on predictive performance than either the smoothing distribution or the proxy-selection criterion. In datasets where a substantial proportion of observations contained missing predictor values, the proposed methods frequently outperformed CART, while maintaining the interpretability and flexibility of tree-based models.

T. S. Prass, A. Neimaier, G. Pumi · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.