Skip to content

Weakly-supervised Learning with Partial Multi-Labels by Leveraging Dual Label Correlation Perspectives

Aug 2026 · ACM Transactions on Knowledge Discovery from Data · 0 citations · 65 references

TL;DR

A novel PML method, namely Wasserstein Partial Multi-Label Learning with dual Label Correlation Perspectives (Wpml3cp), solved by the gradient descent with an augmented Lagrange multiplier technique, and empirical results demonstrate that Wpml3cp and Wpml3cp-D can outperform the PML baselines in various noisy levels.

Abstract

Multi-Label Learning (MLL) refers to inducing multi-label prediction models from the precisely labeled training dataset. However, in many real-world scenarios, e.g., crowdsourcing annotations, the training datasets are often only partially valid, where each training instance is associated with a candidate label set, covering ground-truth labels but also with irrelevant ones. Naturally, learning with such datasets, formally referred to as Partial Multi-label Learning (PML), involves many noisy supervised signals, hence imposing a significant challenge to the prediction model induction. To meet this challenge, we purify the noisy supervised signals by formulating the latent label distribution, i.e., the probability of a candidate label being a ground-truth one, and then jointly learn it with the prediction model by minimizing their regularized Wasserstein distance, i.e., a robust distance for distributions as well as involving label correlations. Therefore, we propose a novel PML method, namely Wasserstein Partial Multi-Label Learning with dual Label Correlation Perspectives (Wpml3cp), solved by the gradient descent with an augmented Lagrange multiplier technique. To further enhance the robustness of Wpml3cp against exceptionally high ratios of irrelevant labels, we extend it with a Dual-branch Competitive Cleansing mechanism, leading to Wpml3cp-D. Besides, we also analyze the generalization error bound and time complexity of Wpml3cp and Wpml3cp-D. The extensive experiments are constructed by comparing Wpml3cp and Wpml3cp-D with existing PML baselines across synthetic and real-world datasets, and empirical results demonstrate that Wpml3cp and Wpml3cp-D can outperform the PML baselines in various noisy levels.

View source

Similar papers

Open access 2026

Partial Multi-Label Learning with Missing Labels via Feature-Aware Label Disentanglement

This work proposes an integrated learning paradigm that simultaneously enhances feature compactness and improves robustness against label noise and introduces a feature disentanglement mechanism that isolates reliable label-related feature representations from spurious ones introduced by noisy supervision.

Yuzhi Tao, Anhui Tan · 0 citations
2025

ComRank: Ranking Loss for Multi-Label Complementary Label Learning

This work proposes ComRank, a ranking loss framework for MLCLL, which encourages complementary labels to be ranked lower than non-complementary ones, thereby modeling pairwise label relationships and ensures Bayes consistency under both uniform and biased cases.

Jin Zhu, Yi Gao, Miao Xu et al. · 0 citations
Preprint Aug 2026

PaSta: Noisy Node Classification with Partial Label Learning

This paper proposes a novel Partial label-based Self-training framework (PaSta) that leverages partial label learning technique to overcome the limitations of existing methods and designs a partial label-based classification model with two well-crafted loss functions to guide the model learning at both label and representation spaces.

Yujing Liu, Yixin Liu, Yu Zheng et al. · 0 citations
2025

Theory-Driven Label-Specific Representation for Incomplete Multi-View Multi-Label Learning

Multi-view multi-label learning typically suffers from dual data incompleteness due to limitations in feature storage and annotation costs. The interplay of heterogeneous features, numerous labels, and missing information significantly degrades model performance. To tackle the complex yet highly practical challenges, we propose a Theory-Driven Label-Specific Representation (TDLSR) framework. Through constructing the view-specific sample topology and prototype association graph, we develop the proximity-aware imputation mechanism, while deriving class representatives that capture the label correlation semantics. To obtain semantically distinct view representations, we introduce principles of information shift, interaction and orthogonality, which promotes the disentanglement of representation information, and mitigates message distortion and redundancy. Besides, label-semantic-guided feature learning is employed to identify the discriminative shared and specific representations and refine the label preference across views. Moreover, we theoretically investigate the characteristics of representation learning and the generalization performance. Finally, extensive experiments on public datasets and real-world applications validate the effectiveness of TDLSR.

Quanjiang Li, Tianxiang Xu, Tingjin Luo et al. · 2 citations
Open access Jul 2026

Graph-Regularized Low-Rank Label Correlation Learning with Label-Specific Features for Missing Labels

Missing labels are common in multi-label learning and can bias both label-correlation estimation and classifier induction. Existing missing-label methods often recover incomplete supervision mainly through global label correlations. However, correlation-driven recovery alone may produce over-smoothed supervision when annotations are sparse, while label-specific discriminative evidence may be weakened. To address this problem, we propose GLCS, a graph-regularized low-rank correlation learning framework with label-specific features for multi-label learning with missing labels. GLCS first uses the observed entries as reliable supervision sources and propagates them through a learned label correlation matrix. It then jointly learns sparse label-specific predictors, low-rank label correlations, and a label graph regularizer induced by the learned correlations. In this way, global label dependencies, local label-structure consistency, and label-wise discriminative features are optimized in a unified objective. The resulting problem is solved by an alternating proximal optimization scheme with soft thresholding for sparse predictors and singular value thresholding for low-rank correlations. Experiments on twelve benchmark datasets under three missing-label ratios show that GLCS obtains strong average performance across AP, AUC, CV, HL, OE, and RL, especially under high missing rates.

Tian-Lin Li, M. F. Nasrudin, Xing-Lin Peng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.