Skip to content
Preprint

Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents

Aug 2026 · 0 citations · 28 references
Computer Science Mathematics

TL;DR

The myopic escalation threshold is derived in closed form, characterise the optimal policy via dynamic programming, and it is proved that the optimal policy is a time-varying threshold with no shape assumption on the raw signal.

Abstract

Current LLM agent systems decide delegation before reasoning begins (a router picks a model) or after a response is complete (a verifier scores it and may retry). We study a third regime: an agent that recognises, during its own reasoning, that it is unlikely to succeed and transfers control to a stronger model. We formulate intra-generation delegation as a Bayesian optimal-stopping problem over a learned competence posterior -- an online estimate of the agent's eventual task success whose sufficient statistics are learned from labelled trajectories, not read off raw entropy. We derive the myopic escalation threshold in closed form, characterise the optimal policy via dynamic programming, and prove that the optimal policy is a time-varying threshold with no shape assumption on the raw signal. We further prove exponential separation of the oracle belief at the Chernoff-information rate of the signal, a regret bound governed by the calibration of the posterior, and a finite-sample guarantee: with n labelled calibration trajectories the deployed plug-in policy's regret decays as 1/sqrt(n). A controlled simulation study confirms each prediction of the theory, including the predicted 1/sqrt(n) rate. We additionally report a real-model validation on a Qwen2.5-Coder 1.5B->7B code cascade (MBPP, 257 tasks), confirming two of three pre-registered predictions: the escalation frontier dominates post-hoc routing at equal cost, and the cumulative competence belief's discrimination rises over generation.

View source

Similar papers

Preprint Aug 2026

Contracting for LLM Delegation: Moral Hazard in Technology and Effort Choice

We extend the standard Principal-Agent framework to scenarios where the Agent selects from a suite of technologies, each characterized by a distinct cost-capability profile. This framework is increasingly critical in the era of Large Language Models (LLMs), where Agents choose both a model and an associated effort level (e.g., token budget). We model the relationship between output quality and effort as a concave, saturating function, which depends on the Agent's hidden two-dimensional action choice balancing technology selection and effort allocation. We derive the optimal linear contract for the Principal, demonstrating that the Agent's best response is characterized by a threshold reward share that triggers technology switching. Finally, we calibrate our model using open-weight LLM pairings across the MATH and MMLUPro benchmarks. We show that both Principal and Agent, when employing bandit algorithms to navigate this environment, converge to strategies that closely align with our theoretical equilibrium. These results suggest that simple linear contracts can effectively incentivize complex, technology-aware delegation in agentic workflows.

Nanda Kishore Sreenivas, Kate Larson · 0 citations
Review Jul 2026

Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games

An auditable framework is built that maintains an external belief state over hidden roles, logs belief updates and belief-action deviations as structured evidence, and supports a defensive offline improvement loop that reviews bad cases before any strategy change.

Yuanpeng Gao, Jiangyi Yang, Yao Zhao et al. · 0 citations
Preprint Aug 2026

Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making

RATTL targets runtime safety for agents, including LLM-based systems, acting under uncertainty, and proves a Safety Sandwich: the RATTL value lies between the uninformed robust value and the full- knowledge optimum, with a gap that vanishes as the posterior concentrates.

D. Ganguly, Jan Křetinský · 0 citations
Jul 2026

When Is Delegated Play Truthful? Within-Range Regret and the Trilemma of Aligned Delegation

Advertisers delegate bidding to autobidders; users delegate tasks to language-model agents. A person describes what they want to an automated proxy that acts in a mechanism on their behalf. This is the revelation principle in production, and it forces a question classical theory assumes away: when is it optimal to describe yourself honestly to your own proxy? We show the answer turns on one quantity, the proxy's within-range regret. The most a principal can gain by misreporting equals the regret of the proxy's honest-report action against those the principal could have steered it to take. Honest self-description is optimal exactly when the proxy already plays the best action it can reach, that is, when it is loyal (Theorem 1). The identity unifies auction-specific autobidding results and pins down when the faithful-communication assumption behind language-model elicitation proxies (Huang et al.) holds. The identity constrains guardrails placed on proxies, from bid caps to a model's alignment layer. No guardrail can be at once binding (it displaces the truthful action from the proxy's best reachable outcome), truthful (honest reporting stays optimal), and capability-preserving (that outcome stays reachable through some report); any two preclude the third (Theorem 2). A safety constraint that alters what a model does while leaving its best output reachable makes honest description of intent suboptimal, so a sharper report can gain. This is the incentive behind prompt-engineering and jailbreaking. Because within-range regret is #P-hard to compute exactly, we estimate it from samples and maintain it as a model is updated, at a cost set by how far the model drifts, not how often it changes. Running it on production language models from five providers under an alignment-style cap, we find honest reporting leaves surplus unclaimed on every model, recovered by inflating the report.

Taksch Dube · 0 citations
Preprint Aug 2026

Random Cap: Optimal Informationally Robust Delegation

Are simple delegation rules optimal under ambiguity? We study delegation when the principal knows the mean, but not the distribution, of the agent's private information. In a parsimonious quadratic constant-bias environment, the robustly optimal randomized mechanism is a random cap: the principal draws and reveals an upper bound, below which the agent chooses freely. Randomization strictly outperforms every deterministic cap by hedging against cap-specific worst-case distributions. We characterize random caps through a nondecreasing and concave expected-action rule and construct the solution using a saddle-point approach. The worst-case distribution features an exponential survival function over its continuous region and an atom at the upper endpoint. Under regularity conditions, the result extends to convex-order ambiguity. Moreover, when the mean is below the agent's bias, an optimum can be implemented by supplementing the random cap with an incentive-neutral outcome lottery, while pure random caps are strictly suboptimal.

Zhiyuan Jia · 0 citations
Preprint Aug 2026

Certified Learning and Equilibrium Implementation under Opaque Partial Commitment

As an extension of existing Bayesian persuasion framework with inadequate message mechanism, we study direct recommendation when a sender is bound by an installed information policy only with probability $\rho$, the realization of binding is hidden, and the receiver does not observe the persistent structural environment. The receiver first sees a payoff-neutral, nonmanipulable calibration sample and then faces a fresh, non-certified deployment interaction. In common, the calibration law identifies only the receiver-facing reduced form, not the latent binding and discretionary kernels. We characterize type-wise $\rho$-implementability, construct the receiver's posterior over the full deployment node, and prove a static direct-following implementation theorem. After every calibration history that passes a posterior-predictive obedience test, the deployment assessment is an exact perfect Bayesian equilibrium: Bayes consistency, receiver sequential rationality, sender sequential rationality, and off-path completion are all verified. Under finite-type separation, common recommendation support, and a positive obedience margin, the test activates such an equilibrium with high probability. Our results keep statistical failure probability distinct from equilibrium approximation. Finally, we embed the original robust value frontier, support-wise linear-programming algorithm, and binary-action fractional-knapsack specialization into this implementation framework

Shuyan Zhang, Xiangtian Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.