Skip to content

Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models

Sep 2026 · 0 citations · 58 references
Computer Science

TL;DR

Pinocchio is introduced, an external calibrator that estimates the correctness of responses from black-box API models that needs only a single forward pass to generate an uncertainty estimate and requires no access to the target model's logits, weights, or internal states.

Abstract

In high-stakes decision-making applications of large language models (LLMs), practitioners require not only accurate LLMs but also uncertainty estimates for their predictions. Existing approaches to uncertainty estimation for LLMs require access to log-probabilities output by the model or require fine-tuning access. However, many industrial LLM products use closed-source API models, and many such API models like GPT do not return log-probabilities and may not allow fine-tuning. We introduce Pinocchio, an external calibrator that estimates the correctness of responses from black-box API models. Trained jointly on responses from seven LLMs, it achieves 0.862 AUROC predicting the correctness of held-out responses from those same models, and shows zero-shot transfer to thirteen unseen models across eight organizations. Our model needs only a single forward pass to generate an uncertainty estimate and requires no access to the target model's logits, weights, or internal states. A lightweight text only 0.8B checkpoint matches our largest model's AUROC. We release code for adding uncertainty estimation to existing repos in only two additional lines of code.

View source

Similar papers

Review Aug 2026

POOL: Propagated Uncertainty Over Lookalikes

This work instantiate this framework with p, a hybrid estimator that combines verbal confidence with spectral answer diversity computed from the negative von Neumann entropy of sampled answer embeddings, and achieves higher average AUROC than verbal confidence and sampling-based uncertainty while using half as many sam...

Rounak Sharma, Ananya B. Sai, Soumyabrata Pal · 0 citations
#natural language process... Preprint Sep 2026

ProbPlug: A Plugin Uncertainty Network for Reliable Confidence in LLM Binary Classification

Experiments show that ProbPlug provides more reliable confidence estimates, improves classification performance with negligible additional overhead, and exhibits strong generalization across tasks, indicating that ProbPlug serves as a practical solution for confidence estimation in LLM-based classification.

Jianzong Wang, Chuhang Liu, Bo-Tao Zhao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

XU-RS: Explaining Credal Width in Random-Set Language Models

Uncertainty estimates tell us how unsure a model is, but not why. Without knowing which parts of an input influences a model's uncertainty, we cannot tell whether that uncertainty score depends on input features that are relevant for the task. We study this problem in randomset classifiers built using pretrained langua...

David Achara, Maryam Sultana, Alexander Rast et al. · 0 citations
Review Aug 2026

Claim-Level Confidence Calibration for Reliable Decision Making with Large Language Models

Claim-level decomposition combined with post-hoc calibration reduces expected calibration error on factual questions while exposing failure modes on adversarial false-premise questions where decision-makers most need reliable uncertainty estimates.

Toghrul Abbasli, Kentaroh Toyoda, Yuan Wang et al. · 0 citations
#natural language process... Preprint Oct 2026

LLM-as-Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them

Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fin...

Yin-Heng Li, Justin Wagle · 1 citation
Preprint Aug 2026

Improved Confidence Estimates for Black-Box Large Language Models

Uncertainty quantification (UQ) is essential for the safe deployment of large language models (LLMs). Existing methods, from verbalized confidence to ones requiring multiple generations, are often zero-shot and produce scores quantifying uncertainty without the need for labelled data. Nonetheless, in practice one must...

S. Mbacke, Mouloud Belbahri, G. Loaiza-Ganem · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.