This work proposes a Bayesian framework that treats clarification as an active learning problem over grounded Signal Temporal Logic task specifications and uses LLMs to initialize candidate formal specifications and translate informative contrasts into natural-language clarification questions, while Bayesian optimization maintains uncertainty estimation over user intent and selects queries that maximize information gain.
Abstract
Interactive robot planning requires robots to infer and execute human intentions from natural language instructions that are often ambiguous, incomplete, or underspecified. Although large language models (LLMs) provide a powerful interface for clarification, relying on the generative model to drive an multi-turn conversation can introduce systematic failures. We propose a Bayesian framework that treats clarification as an active learning problem over grounded Signal Temporal Logic (STL) task specifications. Our method uses LLMs to initialize candidate formal specifications and translate informative contrasts into natural-language clarification questions, while Bayesian optimization maintains uncertainty estimation over user intent and selects queries that maximize information gain. After convergence, the inferred STL specification is passed to a formal planner to synthesize a verifiable robot trajectory. Across four simulated and real-world task domains, our approach generally achieves higher task satisfaction and requires fewer clarification rounds than LLM baselines, while helping smaller models close the performance gap against larger reasoning models.
Closed-Loop contextual Uncertainty rEsolution (Closed-Loop contextual Uncertainty rEsolution), a framework for actively resolving contextual uncertainty given underspecified tasks in natural language, is presented.
Zachary Ravichandran, Jonathan Diller, Fernando Cladera et al.· 0 citations
WorLDS: World-state Observation and Reasoning for Language-guided Discovery and Search, a framework that grounds reasoning in a persistent graph initialized from geospatial priors and updated by perception, is presented.
Finley R. Holt, Luis A. Pabon, J. Alora et al.· 0 citations
This work formalizes the reasoning-execution boundary as a typed contract and constrains language-level decisions through schema-validated tool calls defined by the Model Context Protocol, rejecting malformed commands before they reach the robot.
An active-perception framework for embodied target disambiguation is proposed that uses active observation as the backbone for information acquisition and uses a vision-language model to decide, on the basis of accumulated visual evidence and interaction information, whether to continue observing, request clarification...
Large language models (LLMs) are increasingly used as high-level planners in robot navigation, but their outputs may become unreliable when instructions are ambiguous, unsupported by the environment, or semantically inconsistent. This paper presents a Risk-Aware Semantic Grounding framework for trustworthy LLM-based ro...
Łukasz Sobczak, Nur Keleşoğlu, S. Nowak· 0 citations
BayesBeliefAgent is introduced, which pairs a hierarchical LLM planner with a Bayesian tracking module and evaluates performance using replanning efficiency and the belief-action gap: the fraction of total decisions where an agent with a correct partner estimate executes a non-complementary skill.
Harsh Goel, A. S. Ellendula, Vaishnav Tadiparthi et al.· 1 citation
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.