Skip to content

ExpBoN: Exponential-Noise Best-of-$n$ for Efficient Test-Time LLM Alignment

Sep 2026 · 0 citations · 50 references
Computer Science Mathematics

TL;DR

This paper introduces ExpBoN, an alternative soft BoN method based on the exponential-noise report-noisy-max mechanism that yields exponentially fast convergence in total variation, expected reward, and both directions of KL divergence, and provides comprehensive theoretical analyses of its convergence and regret behavior.

Abstract

Best-of-$n$ (BoN) sampling is a simple yet effective inference-time alignment method, but hard maximization provides only coarse control over the trade-off between reward and distribution shift. Soft Best-of-$n$ (Verdun et al. 2025) provides smoother control and converges to the optimal distribution associated with KL-regularized reward maximization. In this paper, we introduce ExpBoN, an alternative soft BoN method based on the exponential-noise report-noisy-max mechanism. It admits an exact finite-$n$ decomposition, which yields exponentially fast convergence in total variation, expected reward, and both directions of KL divergence. We provide comprehensive theoretical analyses of its convergence and regret behavior. We further integrate ExpBoN into the guided speculative inference (GSI) framework (Geuter, Mroueh, and AlvarezMelis 2025), resulting in ExpGSI, for efficient reward-guided LLM alignment. ExpGSI yields substantial reductions in computational cost while maintaining comparable accuracy. Experiments on MATH500, MMLU-STEM, and Minerva Math with the Qwen2.5-Math and Qwen3 model families show that ExpGSI reduces estimated computation by $14\%$-$39\%$ across candidate budgets for Qwen2.5-Math and by up to $45\%$ at $n=16$ for Qwen3. Overall, our results provide a theoretical and algorithmic foundation for exponential-noise BoN and efficient test-time LLM alignment.

View source

Similar papers

Preprint Aug 2026

Privacy Without Regret: Differentially Private Inference-Time Alignment

Private Inference-Time Pessimism (PrivITP) is introduced, which combines $\chi^2$-regularized rejection sampling with a two-phase Gaussian mechanism, and achieves ex-post $(\epsilon,\delta)$-DP with a privacy cost independent of the number of responses, cleanly decouples the regularization parameter from the privacy pa...

I. Jain, Nandini Bhattad, Sayak Ray Chowdhury · 0 citations
#machine learning Preprint Sep 2026

Minimax-Optimality of Posterior Sampling for Reinforcement Learning

Posterior sampling for reinforcement learning (PSRL) is one of the simplest and most effective exploration methods, but a basic question has remained open: does unmodified PSRL achieve minimax regret without structural assumptions on the prior? We answer yes. Exact vanilla PSRL is minimax optimal in leading-order Bayes...

T. Goo, Kihyuk Hong · 0 citations
Preprint Oct 2026

Best-of-$N$ Guidance for Test-time Diffusion Alignment

Diffusion models achieve strong generative performance but often struggle to align generated samples with human preferences measured by a reward model. A simple yet effective algorithm for test-time alignment is Best-of-$N$ (BoN) sampling, which draws $N$ i.i.d. samples from a pre-trained diffusion model and outputs th...

Richard Kim, Yeongmin Kim, Gyuwon Sim et al. · 0 citations
#machine learning Preprint Oct 2026

Metropolis-Hastings Dominates Importance Resampling for Policy Composition

Post-training a large language model (LLM) often requires exploring trade-offs between multiple rewards, but retraining for each trade-off is expensive. Decoding-time policy composition allows these trade-offs to be adjusted by combining reward-specific policies at inference time. This composition targets a weighted pr...

A. Kurennoy, R. Yarullin, Fergal Reid · 0 citations
Preprint Sep 2026

Accelerated Stochastic Method under $(H_0,H_1)$-Smoothness and Heavy-Tailed Noise

We develop an accelerated stochastic method for convex $(H_0,H_1)$-smooth optimization under heavy-tailed noise. The unbiased oracle has a finite $p$-th noise moment, $1<p\le2$, with constant, gradient-dependent, and gap-dependent terms. Our method combines accelerated updates with clipping, projection, and phase resta...

A. Lobanov, D. Dvinskikh, A. Gasnikov · 0 citations
#natural language process... Preprint Oct 2026

Holdout Best-of-N: Unbiased Evaluation and Its Cost

Reusing the scores that select a Best-of-$N$ winner can overstate its expected reward. We study evaluation from a fixed matrix of $K$ independent scores per candidate for a policy that selects using $J$ fresh scores. A single estimator based only on this matrix is exactly unbiased for expected judge reward under every...

Shrey Shah, Yin-Heng Li · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.