Skip to content

LWCal: Loss-Weighted Calibration for Tabular Classifiers with Noisy Calibration Labels

Sep 2026 · 0 citations · 15 references
Computer Science

TL;DR

LWCal is proposed, a CPU-only post-hoc calibrator that down-weights calibration examples whose noisy labels are contradicted by the base model's held-out probability.

Abstract

Post-hoc probability calibration is usually evaluated under an optimistic assumption: the held-out calibration labels are clean. In many AI deployment settings, however, labels come from weak annotators, historical decisions, heuristics, or distant supervision, so the same label noise that corrupts training also corrupts calibration. We study this overlooked failure mode for tabular classifiers and propose LWCal, a CPU-only post-hoc calibrator that down-weights calibration examples whose noisy labels are contradicted by the base model's held-out probability. LWCal requires no clean validation labels, no noise-rate estimate, and no retraining of the base classifier. A second variant, Gated-LWCal, adds a conservative disagreement gate that backs off toward the raw score when the calibration split appears extremely inconsistent. On nine local binary tabular tasks, six random seeds, symmetric and asymmetric label corruption, and three tree-based base learners, LWCal obtains the lowest average calibration error while Gated-LWCal obtains the best average proper-score tradeoff. In the main random-forest study over 432 noisy cells, Gated-LWCal reduces expected calibration error from 0.188 to 0.122 and negative log likelihood from 0.438 to 0.396 relative to the raw classifier. Paired bootstrap intervals for Gated-LWCal versus raw, Platt, isotonic, and beta calibration exclude zero on ECE, Brier score, and NLL. The artifact contains all scripts, result tables, figures, and the compiled paper.

View source

Similar papers

#machine learning Preprint Sep 2026

SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and Corroboration

SupportCal is introduced, a label-free post-hoc method that retains agreement examples at unit weight and assigns disagreement examples continuous weights based on the own-base PLM's relative support and corroboration from pretrained references selected from a size-compatible candidate pool.

Linhan Luo, Le-Quan Lin, Dai Shi et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Forget or Fine-tune? A Comparative Study of Machine Unlearning Strategies for Noisy Label Correction

This work conducts a comparative empirical study of five MU methods across symmetric, asymmetric, instance-dependent, and open-set noise on CIFAR-10, CIFAR-100, and the real-world noisy dataset Food-101N and finds that the appropriate unlearning strategy is conditioned on the noise structure.

J. L. Sant'Ana, Filipe R. Cordeiro · 0 citations
#machine learning Review Sep 2026

CORDIAL: Calibrating Ordinal LLM Outputs from Few Labels

A large language model (LLM) can turn a text into a distribution over an ordered scale, but that distribution is a noisy measurement: saturated, compressed or exaggerated, and biased in a consistent direction. We propose CORDIAL, which treats the model's output as a noisy reading of the true label and corrects it with...

Xiang-Wei Wang, Peng Wang, Saman K. Halgamuge · 0 citations
Aug 2026

COVER-EL: Class cOVERage-Aware Noise Correction for Crowdsourced Labels via Elimination-Based Inference.

In crowdsourcing scenarios, where each instance is labeled with multiple noisy labels, its true label is estimated by combining label integration and various recently proposed noise correction methods. Recent correction methods typically partition the data into clean and noisy sets, then train classifiers on a high-con...

Bi Wang, Yan-Ting Yang, Xuelian Li · 0 citations
Preprint Aug 2026

C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination

C-Score, a compact framework that evaluates training behavior in three complementary spaces: prediction, feature representation, and optimization, suggests that clean accuracy alone is insufficient for evaluating SSL robustness in open-world environments, and that internal diagnostic signals are necessary for more reli...

Tsao-Lun Chen, Chicheng Fu, Han-Yi Chou et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.