Skip to content

Estimating Rare Events in Language Models with Proper Evaluation

Jul 2026 · arXiv.org · Vol abs/2607.18454 · 0 citations · 37 references
Computer Science

TL;DR

This work introduces Gradient Activation Adaptive Multi-Level Splitting (GA-AMLS), which adapts rare-event Monte Carlo methods to the continuous activation space of language models and establishes activation space as a tractable domain for rare-event estimation in language models, circumventing the brittleness of discrete input-space search.

Abstract

Quantifying the risk of rare failures in language models, such as those triggered by adversarial distribution shifts or very large-scale deployments, requires estimating probabilities far too small for random sampling. While recent work has formalized Low Probability Estimation, existing pipelines remain fragile in the rarest regimes: estimators can suffer zero-estimate collapse or systematic bias, and standard evaluation losses can become unstable or poorly matched to asymmetric safety costs. In this work, we introduce Gradient Activation Adaptive Multi-Level Splitting (GA-AMLS), which adapts rare-event Monte Carlo methods to the continuous activation space of language models. Specifically, GA-AMLS uses a gradient-based MCMC kernel to navigate activation space, eliminating the zero-estimate collapse of input-space search and replacing the independence assumptions of prior activation-space estimators with conditional sampling under an explicit, heavier-tailed activation prior. We also propose the Shifted-Power Bregman (SPB) Loss, a proper scoring rule that remains finite for zero-estimates and offers tunable asymmetry between underestimation and overestimation penalties. Experiments on small transformer models reveal a bias-variance tradeoff: GA-AMLS achieves the lowest loss under symmetric evaluation, reducing average log-space squared error relative to the strongest baseline across model sizes, while methods with overestimation bias prevail under asymmetric penalties. Our findings highlight that estimator choice should be matched to deployment context. More broadly, our work establishes activation space as a tractable domain for rare-event estimation in language models, circumventing the brittleness of discrete input-space search.

View source

Similar papers

Preprint Aug 2026

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

This work proposes the first completely probabilistic architecture-agnostic guardrail to leverage the LLM early output distributional signals for estimating and calibrating the safety probability, thereby enabling early stopping of unsafe ongoing outputs.

Xinzhe Huang, Biwu Yao, Kedong Xiu et al. · 0 citations
Review Aug 2026

POOL: Propagated Uncertainty Over Lookalikes

This work instantiate this framework with p, a hybrid estimator that combines verbal confidence with spectral answer diversity computed from the negative von Neumann entropy of sampled answer embeddings, and achieves higher average AUROC than verbal confidence and sampling-based uncertainty while using half as many samples as Vn10 sampling.

Rounak Sharma, Ananya B. Sai, Soumyabrata Pal · 0 citations
Jul 2026

Diachronic Sample Integration: Robust Tail-Risk Estimation with Generative Models

Diachronic Sample Integration is introduced, a test-time inference framework that ensembles generated samples across checkpoints from a stochastic training trajectory that substantially reduces tail-estimation error compared to single-checkpoint baselines under fixed simulation budgets.

Shuning Zhao, Patrick Wong, Leran Zhang et al. · 0 citations
Jul 2026

Sound Probabilistic Safety Bounds for Large Language Models

We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new application of the Clopper-Pearson confidence intervals to obtain probably approximately correct (PAC) bounds for this problem. As our main technical contribution, we propose an algorithm that leverages features in the latent space to prioritize exploring branches in the auto-regressive generation tree that are more likely to produce harmful outputs. Our approach in particular enables the efficient computation of useful lower bounds, even in scenarios where the true harm probability is extremely small, and crucially, the obtained lower bounds are sound, i.e., formally proven to be less than the actual harmfulness probability: our experimental results demonstrate the effectiveness of our method by computing non-trivial lower bounds on state-of-the-art LLMs. This study newly enables the evaluation and statistical certification of LLMs.

Mahdi Nazeri, Anne-Kathrin Schmuck, S. Soudjani et al. · 0 citations
Jul 2026

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

It is suggested that models possess relevant subpopulation knowledge but do not reliably propagate it into aggregate estimates, and this gap establishes statistical self-consistency as an unsaturated, reference-free criterion for evaluating LLMs.

Patrik Wolf, Thomas Kleine Buening, Andreas Krause et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.