Skip to content
Preprint

PaSta: Noisy Node Classification with Partial Label Learning

Aug 2026 · 0 citations · 51 references
Computer Science

TL;DR

This paper proposes a novel Partial label-based Self-training framework (PaSta) that leverages partial label learning technique to overcome the limitations of existing methods and designs a partial label-based classification model with two well-crafted loss functions to guide the model learning at both label and representation spaces.

Abstract

Noisy node classification problem is a fundamental yet challenging task for real-world graph-related web services, where node labels are often corrupted or unreliable due to weak supervision or automatic annotation. However, existing methods typically train models based on one-hot labels, which not only makes models susceptible to overfitting on noisy labels, but also leads to error accumulation after pseudo-label-guided enhancement. In this paper, we propose a novel Partial label-based Self-training framework (PaSta for short) that leverages partial label learning technique to overcome the limitations of existing methods. Specifically, PaSta first trains multiple annotators to comprehensively capture the class distribution of nodes and aggregates their predictions to construct high-quality partial labels. Subsequently, we design a partial label-based classification model with two well-crafted loss functions to guide the model learning at both label and representation spaces. To further enhance the robustness against noisy labels, we introduce a self-training strategy where the labels refined by partial label learning are then used to further optimize the annotators in a closed-loop iterative manner. Extensive experiments on five datasets demonstrate that, compared with existing state-of-the-art methods, PaSta achieves an average improvement of 1.1% in classification performance under various noise settings.

View source

Similar papers

Open access 2026

Partial Multi-Label Learning with Missing Labels via Feature-Aware Label Disentanglement

This work proposes an integrated learning paradigm that simultaneously enhances feature compactness and improves robustness against label noise and introduces a feature disentanglement mechanism that isolates reliable label-related feature representations from spurious ones introduced by noisy supervision.

Yuzhi Tao, Anhui Tan · 0 citations
Aug 2026

Weakly-supervised Learning with Partial Multi-Labels by Leveraging Dual Label Correlation Perspectives

A novel PML method, namely Wasserstein Partial Multi-Label Learning with dual Label Correlation Perspectives (Wpml3cp), solved by the gradient descent with an augmented Lagrange multiplier technique, and empirical results demonstrate that Wpml3cp and Wpml3cp-D can outperform the PML baselines in various noisy levels.

Ximing Li, Yuanchao Dai, Bing Wang et al. · 0 citations
Open access Aug 2026

Label space reduction for transductive zero-shot classification with large language models

This work proposes distilling the model into a probabilistic classifier, enabling lightweight deployment without repeated LLM calls, and demonstrates that LSR improves macro-F1 scores by an average of 7.0% compared to standard zero-shot classification baselines.

Nathan Vandemoortele, Bram Steenwinckel, F. Ongenae et al. · 0 citations
Open access Jul 2026

CANNE: CLIP-Based ANNE Selection for Noisy-Label Learning

Experimental results on CIFAR-10, CIFAR-100, Animal-10N, and Mini-WebVision, together with additional evaluation under open-set noise, show that the proposed CANNE method achieves competitive performance across diverse noisy-label settings.

Ge Jin, Qian Zhang, Li Huang et al. · 0 citations
Preprint Aug 2026

PseudoMapLabeler: Confidence-Aware Pseudo-Label Generation for Semi-Supervised Online Mapping

This work proposes a teacher-student semi-supervised learning framework that generates high-quality pseudo-labels from unlabeled data through confidence-aware map refinement, and introduces a spatial clipping technique that selectively preserves high-confidence regions while removing unreliable segments.

Chikao Tsuchiya, Dhaval Bhanderi, David Ilstrup et al. · 0 citations
Preprint Aug 2026

Dual-Primal Graph VAEs for Noisy Label Aggregation

Inferring the ground-truth from noisy crowdsourced labels is an important theoretical and practical problem. Neural network-based methods offer an alternative to classical Bayesian models which require specifying a family of generative models used for inference. However, current models either still rely on fairly simple generative models for inference or require pseudo-labels or synthetic data to train the aggregate classifier. We propose a graph VAE architecture in which the decoder and encoder use GAT-based message passing on the adjacency graph of a crowdsourced dataset and its dual, respectively. The ground-truth labels are treated as latent variables, enabling unsupervised representation learning without needing to train a separate classifier. We show our model achieves state of the art performance on crowdsourcing benchmarks. We then demonstrate the generality of our approach by showing how the original crowdsourcing graph can be augmented to incorporate side information such as representations from neural network classifiers trained on the noisy labels to substantially boost their classification performance at test time.

Patrick Stinson, N. Kriegeskorte · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.