Heterogeneous compositional data may be simultaneously affected by missing values and atypical points, posing challenges for both clustering and outlier detection. We develop a mixture model for incomplete compositional data under Huber's contamination model, with contamination defined directly on the simplex and on observations that may be missing at random. The model provides a principled representation of outliers and allows the distribution of missing parts to be derived while accounting for contamination. We establish that maximisation of the observed log-likelihood constructed from contaminated, mean-parametrised Dirichlet densities is a convex optimisation problem. We then develop a tailored expectation-maximisation. The E-step incorporates the moments from the distribution of the missing parts of the data. Although the resulting parameter estimates are not available in closed form, the maximisation step admits tractable element-wise iterative updates. Numerical experiments demonstrate the performance of the proposed approach under varying percentage of missingness and contamination, and different sample size. An application to the American Time Use Survey identifies two interpretable clusters corresponding to work-intensive and sociable recreational days, while revealing atypical time-use compositions. In contrast, a conventional Dirichlet mixture model identifies four clusters, reflecting the influence of outliers and an artificial splitting of one cluster.
J. Pillay, A. Bekker, C. Tortora et al.· 0 citations
Likelihood-based inference for compositional data generally requires fully observed compositions, hindering the direct treatment of missing or censored components on the simplex. In this paper, we develop an expectation-maximisation (EM)-type algorithm for maximum likelihood estimation of the Dirichlet parameters in the presence of missing and censored components under a unified coarsening framework. The Dirichlet distribution---the canonical probability model for compositional data, which plays a role analogous to that of the multivariate normal distribution for unconstrained multivariate data---provides the foundation for our methodology. Our methodology preserves the compositional structure of the data while simultaneously performing parameter estimation and model-based imputation. We evaluate the performance of our estimators and imputations through a simulation study under increasingly complex coarsening mechanisms, including both missing and censored data. We compare our method with an existing model-based approach and a nonparametric alternative. Finally, we illustrate the practical utility of our methodology using mercury speciation data, in which compositions are only partially observed because of detection limits and incomplete speciation. Our results indicate that the Dirichlet distribution provides a suitable model for these data and that our method yields imputations that better preserve the observed compositional structure than competing approaches.
J. Pillay, A. Bekker, C. Tortora et al.· 1 citation
Incomplete compositional data analysis faces a fundamental limitation: likelihood-based methods for compositional models generally require fully observed compositions, making it difficult to accommodate missing or censored proportions directly on the simplex. Consequently, analysts often discard partially observed compositions or transform the data into unconstrained spaces, potentially sacrificing interpretability and coherence. This paper proposes a likelihood-based method for incomplete compositional data without leaving the simplex. Specifically, we develop an Expectation-Maximisation (EM) type algorithm for fitting finite mixtures of Dirichlet distributions in the presence of missing and censored components. The proposed approach performs parameter estimation and model-based imputation simultaneously while preserving the compositional structure and interpretability of the original variables. A simulation experiment evaluates the performance of the proposed estimators and imputations under increasingly complex coarsening mechanisms. Particular attention is paid to clustering performance, and model selection outcomes. The results showed beneficial clustering performance despite observations being incomplete, and a higher probability of model selection metrics identifying the correct number of clusters compared to current alternative of case-deletion. The practical utility of the method is illustrated using two real datasets with distinct coarsened patterns. Analysis of the xenolith dataset identifies a four-component Dirichlet mixture that reveals interpretable profiles of rock types and speciation methods. Application to PM$_{2.5}$ speciation data from the Air Quality System, containing both left-censored and missing-at-random values, supports a four-component mixture model that characterises compositional parts of particulate matter across the United States.
J. Pillay, A. Bekker, C. Tortora et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.