Skip to content
Preprint

hint$^2$: Hierarchical World Models for Inference-Time Temporal Logic Guidance

Aug 2026 · 0 citations · 58 references
Computer Science

TL;DR

This paper introduces hint, a method for guiding short-horizon policies toward satisfying complex LTL specifications at inference time using hierarchical world models, and shows that hint$^2$ can handle complex instructions on a real UR5e manipulator.

Abstract

A central goal of robot learning is to enable robots to execute rich instructions specified at runtime. Large-scale language-conditioned policies have made substantial progress toward this goal, yet still struggle with temporal structure and safety constraints. Linear Temporal Logic (LTL) provides a powerful language to express complex, non-Markovian instructions. However, guiding learned manipulation policies toward LTL satisfaction remains challenging because modern policies generate short-horizon action chunks and replan in closed loop, while almost all LTL specifications are evaluated over long-horizon trajectories. In this paper, we introduce hint$^2$, a method for guiding short-horizon policies toward satisfying complex LTL specifications at inference time using hierarchical world models. Our key idea is to derive two separate guidance objectives using each world model's abstraction level. A high-level model predicts future action-induced transitions in task-relevant atomic propositions to guide progress through the LTL automaton, while a low-level dynamics model predicts immediate state evolution for accurate local safety guidance. Our results show that hint$^2$ overcomes the limitations of current LTL-guided diffusion methods, outperforms existing inference-time steering methods in CALVIN, and successfully completes instructions with complex liveness and safety constraints more elegantly than language-conditioned alternatives. Finally, we demonstrate that hint$^2$ can handle complex instructions on a real UR5e manipulator.

View source

Similar papers

Jul 2026

STeP: Signal Temporal Logic for Precise Specifications for Action Generation with Vision Language Models

A hierarchical framework that uses Signal Temporal Logic (STL) as a shared representation connecting high-level language understanding with low-level robot execution is proposed, demonstrating how formal specifications can improve the precision, reliability, and interpretability of language-conditioned robot planning.

Kasra Torshizi, Anukriti Singh, Sidharth Mathur et al. · 1 citation
Jul 2026

A Few Words Go a Long Way: Language Guided Robot Policy Synthesis

The results demonstrate that the synthesized skill library enables the system to transfer to novel tasks with decreasing human intervention, providing a steerable and data-efficient alternative to black-box robot learning.

Daphne Chen, A. Jain, E. Goossen et al. · 0 citations
Jul 2026

SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning

Experiments show that coupling latent subgoal decomposition with prior-conditioned action generation substantially improves long-horizon planning while preserving strong short-horizon performance.

Le-Tian Cheng, Qi Zhang, Yisen Wang · 2 citations
Review Open access Jul 2026

Large Language Models for Task Planning in Embodied AI: A Survey

A structured taxonomy is presented that organizes existing work into three complementary paradigms that represent dominant architectural tendencies in current LLM-based embodied task planning research, and compares these paradigms along dimension of accuracy, robustness, scalability, efficiency, and sim-to-real transfer.

Zhen Zhang · 0 citations
Preprint Aug 2026

$R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning

This paper introduces a simple post-training recipe that turns off-the-shelf VLMs into robotic reasoners, and suggests that free-form language reasoning can function as a test-time compute mechanism for steering low-level policies.

Lehong Wu, Yuxiao Qu, Zheyuan Hu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.