Skip to content
Review Open access

Conformal prediction for multi-label learning: a review of methods and guarantees.

Aug 2026 · Philosophical transactions. Series A, Mathematical, physical, and engineering sciences · Vol 384 2327 · 1 citation · 27 references
Medicine

TL;DR

This review consolidates the landscape of CP adaptations for MLL under a unified framework, examining the types of outputs and guarantees they provide, where label dependencies are incorporated, and how inference cost scales with the number of labels.

Abstract

Multi-label learning (MLL) is a machine learning paradigm that aims to predict a set of labels for each instance, rather than a single class. Such tasks arise in a wide range of real-world applications and pose significant challenges, including an exponentially large output space, dependence among labels and often severe label imbalance. These challenges amplify predictive uncertainty, making reliable uncertainty quantification essential. Conformal prediction (CP) is an attractive answer: it converts model outputs into prediction regions with distribution-free, finite-sample guarantees under the sole assumption of data exchangeability. Several adaptations of CP to the multi-label setting have been proposed. Yet these vary widely in scoring constructions, output types and targeted guarantees. This review consolidates the landscape of CP adaptations for MLL. It places existing approaches under a unified framework, examining the types of outputs and guarantees they provide, where label dependencies are incorporated, and how inference cost scales with the number of labels. It provides an in-depth analysis of all approaches using common notation, identifying their key characteristics along with their practical implications and assessing their strengths and limitations. Finally, it compares approaches side-by-side, highlighting trade-offs among guarantee types, precision of regions, compactness of outputs and scalability. This article is part of the theme issue 'Advancing uncertainty quantification in AI systems'.

Read PDF

Similar papers

Preprint Aug 2026

Online Conformal Prediction Beyond Feedback

Uncertainty quantification is essential when deploying machine learning models in safety-critical applications. Online conformal prediction (OCP) provides theoretically principled uncertainty quantification for arbitrary black-box classifiers and non-i.i.d. data streams by constructing prediction sets that are guaranteed to contain the true label at a user-specified frequency. OCP usually updates prediction sets using feedback from previously deployed predictions. We instead study an OCP setting beyond feedback: on each round, the learner can either output a prediction set or query the correct label, but not both. Thus, no deployed prediction is ever evaluated directly. We reduce this problem to a partial monitoring game in which prediction actions return no observation and a separate query action reveals the label. The reward function is constructed in a way that encourages the learner to output small prediction sets while ensuring that the correct label is covered with a sufficiently high probability. To solve this game, we develop OCP with queries (OCPQ) by adapting the label efficient forecaster of Cesa-Bianchi, Lugosi, and Stoltz (2004) to our setting. For any black box classifier and any (non-i.i.d.) oblivious data stream of length $T$, OCPQ has $O(T^{2/3})$ expected regret and expected coverage at least $\beta-O(T^{-1/3})$ for a user-defined $\beta$, while querying only an expected $T^{-1/3}$ fraction of rounds. This provides coverage comparable to bandit-based OCP methods while requiring no feedback from deployed prediction sets. Experiments on real-world datasets further demonstrate the effectiveness of our approach.

J. Skalse, Edoardo Pona, Osvaldo Simeone et al. · 0 citations
Aug 2026

Multi-label classification with extreme learning machine and twin support vector machine: a novel hybrid framework

This work proposes a novel hybrid approach that combines the strengths of the extreme learning machine (ELM) and the twin support vector machine (TSVM) to address the challenges of robustness and scalability in multi-label classification, particularly in settings where deep learning is not practical due to limited training instances.

Amisha Bharti, Vasudha Bhatnagar, Vikas Kumar · 0 citations
2025

ComRank: Ranking Loss for Multi-Label Complementary Label Learning

This work proposes ComRank, a ranking loss framework for MLCLL, which encourages complementary labels to be ranked lower than non-complementary ones, thereby modeling pairwise label relationships and ensures Bayes consistency under both uniform and biased cases.

Jin Zhu, Yi Gao, Miao Xu et al. · 0 citations
Aug 2026

Weakly-supervised Learning with Partial Multi-Labels by Leveraging Dual Label Correlation Perspectives

A novel PML method, namely Wasserstein Partial Multi-Label Learning with dual Label Correlation Perspectives (Wpml3cp), solved by the gradient descent with an augmented Lagrange multiplier technique, and empirical results demonstrate that Wpml3cp and Wpml3cp-D can outperform the PML baselines in various noisy levels.

Ximing Li, Yuanchao Dai, Bing Wang et al. · 0 citations
Open access Aug 2026

Inductive Conformal Prediction for Guaranteed Class-Label Coverage in Object Detection

Conformal prediction has emerged as a principled framework for uncertainty quantification in computer vision, offering rigorous finite-sample coverage guarantees. However, its application in object detection has remained largely confined to localization, as standard inference codebases typically yield only top-1 class scores, precluding full class-label conformalization. In this work, we bridge this gap by adapting four architecturally diverse detectors—Faster R-CNN, RetinaNet, YOLO11, and RT-DETRv2—to facilitate the extraction of comprehensive per-class score vectors and the estimation of background confidence in the absence of native background modeling. Leveraging these adapted architectures, we implement inductive conformal prediction (ICP) using five distinct nonconformity functions: Top-K, Adaptive Prediction Sets (APS), Hinge, Margin, and Brier score. Our framework is rigorously benchmarked across a curated 20-class subset of MS-COCO and two specialized parasite egg datasets (AI4NTD P1.5v2 and Chula-ParasiteEgg-11). In addition, a Naive cumulative-threshold method is included as a baseline for comparison with APS, given their comparable mathematical formulations. Across target coverage levels of 90%, 95%, and 99%, the conformalized models consistently achieved nominal coverage with only minor finite-sample deviations. Hinge and APS exhibited an optimal balance between statistical coverage and prediction-set efficiency, whereas Margin and Brier scores tended toward larger sets under high data complexity and strict coverage requirements. With empty prediction sets maintained below 0.1%, our findings establish ICP as a robust and adaptable paradigm for trustworthy class-label uncertainty estimation, particularly within safety-critical workflows such as automated parasite diagnostics.

M. A. Mohammed, Esla Timothy Anzaku, Jef Jonkers et al. · 0 citations
Review Open access Jul 2026

Advancing multi-class classification: innovations, challenges, and ethical perspectives in machine learning

This review examines recent advances and persistent challenges in multi-class classification within machine learning (ML) and deep learning (DL), a core task underpinning many real-world applications in healthcare, finance, social media, and other high-impact domains. The review provides a structured analytical synthesis of major methodological directions, including problem transformation methods, algorithm-level approaches, ensemble and hybrid strategies, class-imbalance handling, evaluation metrics, and deployment-related considerations such as interpretability, uncertainty, and ethics. Recent progress in deep learning architectures, ensemble learning, and active learning has substantially improved predictive accuracy, robustness, and data efficiency across diverse application settings. At the same time, important challenges remain, including imbalanced datasets, noisy labels, scalability constraints, computational cost, and the need for trustworthy and transparent decision-making. The review also examines the trade-offs among predictive performance, interpretability, fairness, and deployment feasibility, highlighting the importance of selecting methods and evaluation metrics that are appropriate to application context. Rather than proposing a new algorithm, this work offers an integrative framework that connects technical advances with practical and ethical considerations in multi-class classification. Finally, it identifies key future directions, including semi-supervised and transfer learning, few-shot and federated multi-class systems, robustness under label noise, energy-efficient model design, and scalable interpretability frameworks for high-stakes deployment.

Y. Qawqzeh, Abdullah Alourani, Fayez Alharbi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.