By separating different pseudo-label reliability regions into specialized adaptation paths, TriNoL improves robustness to noisy supervision while keeping the training cost low.
Abstract
Semi-supervised adaptation of vision foundation models (VFMs) commonly freezes the pretrained backbone and updates lightweight modules such as LoRA. However, pseudo-labels have mixed reliability, and a single LoRA adapter must absorb reliable, ambiguous, and noisy gradients in the same low-rank space. This can make VFM adaptation sensitive to pseudo-label noise. We propose \textbf{TriNoL}, a \textbf{Tri}ple-expert learning framework from \textbf{No}isy \textbf{L}abels for semi-supervised VFM adaptation. TriNoL routes unlabeled samples into three confidence regions and assigns them to three LoRA experts: a Positive Expert for high-confidence pseudo-labels, an Alignment Expert for medium-confidence ambiguous samples, and a Negative Expert for low-confidence noisy samples. The VFM backbone remains frozen, and only the LoRA experts and classifier head are updated. By separating different pseudo-label reliability regions into specialized adaptation paths, TriNoL improves robustness to noisy supervision while keeping the training cost low.
Experimental results on CIFAR-10, CIFAR-100, Animal-10N, and Mini-WebVision, together with additional evaluation under open-set noise, show that the proposed CANNE method achieves competitive performance across diverse noisy-label settings.
Ge Jin, Qian Zhang, Li Huang et al.· Entropy· 0 citations
Reliable abnormal behavior recognition from surveillance videos is hindered by the high cost of clip-level annotation, the scarcity of abnormal samples, and the context-dependent nature of behavioral semantics. Although vision–language models offer strong semantic transferability, their application under limited supervision remains susceptible to noisy pseudo-labels and confirmation bias. We propose confidence-aware semi-supervised vision–language contrastive learning (CA-VLC), which jointly exploits limited labeled videos and abundant unlabeled videos. Building on an existing CLIP-initialized temporal backbone, CA-VLC combines behavior-only and context-enriched text prototypes through confidence- and agreement-guided semantic fusion. For unlabeled videos, the model generates predictions from weakly augmented views and selects reliable pseudo-labels using entropy-based confidence estimation and class-adaptive thresholds. Detached weak-view targets then supervise strongly augmented views through confidence-weighted self-training without requiring an additional teacher network. Furthermore, cross-view consistency regularization and confidence-aware contextual alignment suppress unreliable semantic cues and improve robustness to contextual noise. Experiments on CABR50 demonstrate consistent improvements across multiple labeled-data ratios, while evaluations on CABRZ6 and UCF-101 assess prompt-based transfer to predefined target label sets without target-domain fine-tuning. With 10% labeled videos, CA-VLC achieves 84.06% Top-1 accuracy and 83.51% Macro-F1, retaining 95.47% of its fully supervised Top-1 accuracy of 88.05%, thereby demonstrating its effectiveness for label-efficient abnormal behavior recognition.
Results on the UC Merced (UCM) and NWPU benchmarks indicate that SE-CLIP significantly outperforms existing semi-supervised approaches and provides a viable solution for adapting VLMs to the remote sensing domain with minimal human intervention.
M. L. Mekhalfi, M. M. Al Rahhal, Y. Bazi et al.· IEEE Geoscience and Remote S...· 0 citations
This work proposes a teacher-student semi-supervised learning framework that generates high-quality pseudo-labels from unlabeled data through confidence-aware map refinement, and introduces a spatial clipping technique that selectively preserves high-confidence regions while removing unreliable segments.
Chikao Tsuchiya, Dhaval Bhanderi, David Ilstrup et al.· 0 citations
CAPSUN is proposed, a robust framework that improves the precision of clean sample selection and mitigates distribution bias through alignment among subsets, and designs a distribution alignment module to adjust the class distribution contrast of labeled and unlabeled subsets to mitigate class distribution discrepancies.
Weakly supervised semantic segmentation (WSSS) with image-level labels is largely limited by the reliability of dense seed supervision. Existing CAM- and CLIP-based methods provide complementary localization cues, but their predictions are biased in different ways: classification-oriented cues are usually precise but incomplete, whereas alignment-oriented cues offer broader coverage but are more susceptible to contextual noise. In this letter, we propose Reliability-Calibrated Posterior Supervision (RCPS), a simple yet effective framework that uses discriminative classification evidence to calibrate broad CLIP-oriented cues. RCPS first constructs a semantic target through classification-regularized posterior calibration, avoiding direct commitment to either noisy alignment responses or incomplete classification activations. It then estimates pixel-wise reliability from cue self-certainty, classification–alignment consistency, and confidence-preserving disagreement, retaining potentially complementary evidence while suppressing uncertain responses. The resulting reliability-weighted objective learns dense seeds from calibrated soft targets. Experiments on PASCAL VOC 2012 and MS COCO 2014 show that RCPS consistently improves seed quality, pseudo-mask quality, and final segmentation performance over strong CAM- and CLIP-based baselines.
Xiao-Ya Sun, Xin Xu· IEEE Signal Processing Lette...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.