Preprint
Aug 2026
Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents
The myopic escalation threshold is derived in closed form, characterise the optimal policy via dynamic programming, and it is proved that the optimal policy is a time-varying threshold with no shape assumption on the raw signal.
Nadeem Shaikh
· 0 citations