Aug 2026· Frontiers in Robotics and AI· Vol 13· 0 citations· 30 references
Medicine
TL;DR
The results indicate that predictive vision-language monitoring can improve task completion in simulated dynamic robot tasks, while remaining subject to limitations such as VLM latency, prompt sensitivity, and evaluation beyond simulation.
Abstract
Robots that execute language-conditioned tasks in dynamic environments often rely on feedback only after an action has failed, which can be insufficient when failures involve collisions or workspace conflicts. This paper presents a predictive monitoring framework that uses Vision-Language Models (VLMs) to assess near-future execution risk during robot task execution. The framework first generates structured plans with action execution conditions and a plan-level fallback action. During execution, a monitoring module combines visual observations, the current action, and the relevant execution conditions to estimate whether a condition is likely to be violated within a short future time window. When the predicted risk exceeds a task-specific threshold, the robot halts the current action, executes the fallback behavior, and replans from the updated state. We evaluate the approach in Gazebo simulation on mobile navigation with a moving human obstacle and manipulation with two robot arms sharing a workspace. Across controlled collision-risk settings, the proposed method achieves higher task success rates than reactive VLM-based baselines while requiring fewer replanning events than conservative current-state precondition checking. The results indicate that predictive vision-language monitoring can improve task completion in simulated dynamic robot tasks, while remaining subject to limitations such as VLM latency, prompt sensitivity, and evaluation beyond simulation.
Evidence Acquisition and Feasibility Gating (EAFG) is proposed, a framework that acquires visual evidence through VLM-generated exploratory subgoals and TAMP-based execution and applies a feasibility gate to decide whether to proceed with task planning, acquire further evidence, or halt.
Tsunehiko Tanaka, Matthew Stephenson, Alistair Macvicar et al.· 0 citations
ActFovea is introduced, a plug-and-play safeguarding framework that detects and mitigates runtime failures of vision-language-action policies without retraining or modifying the underlying VLA policy.
Wen-Da Yu, Tianshi Wang, Fengling Li et al.· 0 citations
This paper proposes FutureRTC, a plug-and-play adaptation framework that predicts execution-time observations and states for asynchronous VLA control without modifying the underlying policy, and introduces a policy consistency loss to align the action chunks generated from predicted contexts with those produced under the expected execution-time inputs of the VLA policy.
Hai Jiang, Yi Zou, Binbin Liang et al.· arXiv.org· 0 citations
The World-Cognition Model is presented, a human-centered embodied agent built on the SLAK architecture (Sensing, Logic, Action, and Knowledge) and an asynchronous runtime and introduces a human-in-the-loop teaching mode that enables users to interactively teach the robot difficult or long-horizon tasks.
A language-guided robot operating in a real kitchen must do more than produce a plan that appears correct. It must also execute that plan safely in cluttered environments under imperfect perception. Large language models (LLM) can decompose instructions into action sequences, yet a language-action gap remains: a plan may appear valid linguistically while being physically infeasible under kinematic and collision constraints. We bridge this gap by formalizing the reasoning-execution boundary as a typed contract. From RGB-D observations, the system grounds perceived objects in an explicit, collision-aware scene model and constrains language-level decisions through schema-validated tool calls defined by the Model Context Protocol (MCP), rejecting malformed commands before they reach the robot. Each validated call is deterministically grounded in a MoveIt Task Constructor pipeline, where candidate motions are evaluated against the reconstructed planning scene in a verify-then-act step. Only trajectories that pass both kinematic and collision checks are sent to the robot. On a physical UFactory 850, the method achieves up to 80% success across ten trials per task on pouring tasks involving liquids, granular media, and discrete solids. It achieves 90% success on a grasp-and-place task using the same planning, protocol, and verification stack. Although a scripted policy slightly outperforms our method on the easiest task, its success rate falls to 10% on the hardest, compared with 60% for our method.
Results demonstrate that model-external structured state maintenance and closed-loop agentic decision making can effectively extend the local control capabilities of WAMs into embodied task execution that is plannable, verifiable, and recoverable.
Zhaopeng Gu, Bingke Zhu, Tianxin Lin et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.