Skip to content
Preprint

Favourable Missingness in Semi-Supervised Classification for Exponential Mixture Models

Aug 2026 · 0 citations · 20 references
Mathematics

TL;DR

This work studies a different regime in which the probability of label missingness depends on posterior classification uncertainty, so that the observed missing-label indicators can themselves carry information about the Bayes decision boundary.

Abstract

Semi-supervised classifiers are commonly trained from samples in which all features are observed but some class labels are missing. When label missingness is independent of the observed data, unavailable class memberships reduce Fisher information relative to a completely classified sample. We study a different regime in which the probability of label missingness depends on posterior classification uncertainty, so that the observed missing-label indicators can themselves carry information about the Bayes decision boundary. Building on the conditionally weighted information decomposition of Ahfock and McLachlan, we develop this phenomenon for a two-component exponential mixture. Although the exponential model is non-Gaussian, asymmetric, and supported on the positive half-line, its log-posterior odds remain linear in the feature. We derive Bayes'rule and its exact error rate, formulate entropy-logistic and squared-discriminant missingness mechanisms, and obtain the full partially classified likelihood. We then derive a decomposition of the Fisher information into the complete-data information, the conditionally weighted loss due to missing labels, and the information contributed by the missing labels. Numerical quadrature identifies regions in which the full likelihood classifier has asymptotic relative efficiency above or below one. Monte Carlo experiments with finite training samples broadly support the population calculations, with the largest departures from the asymptotic predictions occurring near the transition at which the relative efficiency crosses one.

View source

Similar papers

#machine learning Preprint Sep 2026

Semi-Supervised Classification with Informative Missing Labels in Weibull Mixture Models

We consider semi-supervised classification from a partially classified sample arising from a two-component Weibull mixture. The feature is observed for all data, whereas some class labels are missing. The probability of a missing label is modelled as a function of classification uncertainty, giving a feature-dependent missing-at-random (MAR) mechanism that shares parameters with the Weibull-mixture classifier. The missing-label indicators can therefore provide information about the classifier in addition to the observed features and available class labels. Under a common Weibull shape, a Bayes'rule has at most one positive decision boundary, which is unique when the rule is nonconstant; under unequal shapes, it can have two. We characterise these decision regions, derive the Fisher information for the classifier after adjustment for nuisance parameters in the missingness model, and obtain a decision-boundary expansion of the expected error rate of the plug-in sample rule relative to the Bayes error. The expansion yields classification-specific asymptotic relative efficiency formulas for the one- and two-boundary cases and shows that a positive-definite increase in Fisher information is sufficient, but not necessary, for a smaller first-order expected error rate. Numerical studies and a semi-synthetic analysis based on hard-drive failure data illustrate potential reductions in expected error rate and improvements in decision-boundary estimation from modelling feature-dependent label missingness.

Jinran Wu, You‐Gan Wang, Geoffrey J. McLachlan · 0 citations
#machine learning Preprint Aug 2026

Informative Label Missingness in Multiclass Classification Information Geometry and Excess Risk

A classification-weighted generalized-eigenvalue criterion is developed under which informative partial classification may have smaller asymptotic classification risk without globally dominating complete classification in Fisher information.

Fariborz Setoudehtazang, Geoffrey J. McLachlan · 0 citations
Jul 2026

Multiclass Classification without Labels via Posterior Simplex Geometry

Classification without Labels (CWoLa) shows that, in the binary case, a classifier trained to distinguish two impure mixtures with different class proportions can recover an optimal class discriminator without knowing the mixture proportions, and proposes prior-free procedures that train a standard classifier to distinguish mixture identities and then extract latent class structure using either post-hoc simplex fitting or a bottleneck architecture.

Raphaël Bonnet-Guerrini, Johann Ioannou-Nikolaides, Troels C. Petersen et al. · 1 citation
#machine learning Preprint Sep 2026

Large Classification-Risk-Optional Label Acquisition

We study how a limited labeling budget should be allocated to minimize multiclass zero-one classification risk. We consider parametric classification problems in which features are observed for all sampling units while class labels can be acquired selectively. By combining the Fisher information supplied by an acquired label with the local geometry of multiclass excess risk, we derive an acquisition criterion that minimizes the leading asymptotic coefficient of expected multiclass excess risk. The resulting rule values a label according to how strongly its information is aligned with parameter directions that perturb the active Bayes decision boundary, rather than according to posterior uncertainty or global parameter information alone. We characterize the oracle acquisition design, establish its threshold structure, and derive face-specific and cost-sensitive extensions. An analytic example shows that posterior uncertainty and classification value can produce different, and even reversed, acquisition rankings. We further develop a two-stage adaptive procedure that attains the oracle leading-risk criterion under regularity conditions and provide explicit results for Gaussian discriminant analysis. Three-class QDA experiments illustrate the resulting acquisition geometry, while an application to the six-class Statlog Landsat Satellite data shows that classification-risk acquisition can differ materially from both uncertainty-based acquisition and the complete-classification-information comparator. The adaptive classification-risk design attains lower mean error than this Fisher comparator across the labeling budgets considered, although it does not uniformly outperform entropy or margin sampling and differences among the targeted strategies become small as the labeling budget increases.

F. Setoudehtanzangi, Geoffrey J. McLachlan · 0 citations
Preprint Jul 2026

Handling Missingness and Censoring in Dirichlet Mixture Models

Incomplete compositional data analysis faces a fundamental limitation: likelihood-based methods for compositional models generally require fully observed compositions, making it difficult to accommodate missing or censored proportions directly on the simplex. Consequently, analysts often discard partially observed compositions or transform the data into unconstrained spaces, potentially sacrificing interpretability and coherence. This paper proposes a likelihood-based method for incomplete compositional data without leaving the simplex. Specifically, we develop an Expectation-Maximisation (EM) type algorithm for fitting finite mixtures of Dirichlet distributions in the presence of missing and censored components. The proposed approach performs parameter estimation and model-based imputation simultaneously while preserving the compositional structure and interpretability of the original variables. A simulation experiment evaluates the performance of the proposed estimators and imputations under increasingly complex coarsening mechanisms. Particular attention is paid to clustering performance, and model selection outcomes. The results showed beneficial clustering performance despite observations being incomplete, and a higher probability of model selection metrics identifying the correct number of clusters compared to current alternative of case-deletion. The practical utility of the method is illustrated using two real datasets with distinct coarsened patterns. Analysis of the xenolith dataset identifies a four-component Dirichlet mixture that reveals interpretable profiles of rock types and speciation methods. Application to PM$_{2.5}$ speciation data from the Air Quality System, containing both left-censored and missing-at-random values, supports a four-component mixture model that characterises compositional parts of particulate matter across the United States.

J. Pillay, A. Bekker, C. Tortora et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.