GRAB-TAMP is introduced, an FM-based TAMP framework that searches for scene entities required for task completion, grounds functional roles to valid physical objects, and plans only after a complete joint assignment establishes functional sufficiency.
Abstract
Foundation models (FMs) have expanded task and motion planning (TAMP) to manipulation problems specified through language and visual observations. However, incomplete scene knowledge leaves a critical gap between understanding what the task requires and knowing whether the physical scene can actually realize it. We introduce GRAB-TAMP, an FM-based TAMP framework that searches for scene entities required for task completion, grounds functional roles to valid physical objects, and plans only after a complete joint assignment establishes functional sufficiency. We represent the task through functional roles, relations, and assignment constraints, and incrementally inspect the scene while requirements remain unresolved, verifying candidate objects through semantic, geometric, and relational checks. We evaluate GRAB-TAMP across 32 scene variants spanning Kitchen, Living Room, and Workshop domains. Across 200 feasible trials, our approach achieves 54.0% end-to-end success with 67.3% plan goal coverage. Compared with three FM-based TAMP frameworks under the same execution setting, GRAB-TAMP improves end-to-end success by 25.7 percentage points over the mean baseline. Implementation and evaluation code: https://github.com/Narendhiranv04/GRAB-TAMP
This paper provides a formalization of what constitutes a sufficient scene graph for planning by modeling planning over scene graphs within an information-spaces framework through the definition of scene graph transition systems and relevant action semantics for navigation and manipulation.
Evidence Acquisition and Feasibility Gating (EAFG) is proposed, a framework that acquires visual evidence through VLM-generated exploratory subgoals and TAMP-based execution and applies a feasibility gate to decide whether to proceed with task planning, acquire further evidence, or halt.
Tsunehiko Tanaka, Matthew Stephenson, Alistair Macvicar et al.· 0 citations
Robots operating in cluttered environments must often manipulate objects whose locations are only partially observable. A central challenge is deciding whether to acquire another observation or to first manipulate objects that may occlude the target. Conventional task and motion planning (TAMP) approaches typically mak...
Antareep Singha, Shiv Kumar, Yoonwoo Kim et al.· 0 citations
Robot demonstration generation requires a system to identify where an interaction should occur, plan a feasible motion, and execute the required contact. HiWE connects these decisions through a point-based interface between visual grounding and language-based planning. PointVLM is instruction-tuned to associate task-re...
Guo-Qing Ma, Ming-Qi Yuan, Chen Gao et al.· 0 citations
Experimental results demonstrate that current methods struggle to complete the ESRP task efficiently, highlighting ESRP as a challenging frontier for embodied agents in scene understanding and long-horizon task planning.
Can-Zhi Chen, Zan Wang, Si-Qi Zhu et al.· IEEE Robotics and Automation...· 0 citations
This work proposes a method that leverages a TAMP approach, defining object-centric abstractions of execution constraints, called Unified TAMP (U-TAMP), to execute robotic tasks involving interactions among objects with heterogeneous shapes, sizes, and materials.
Pouya P. Niaz, Justus H. Piater, Alejandro Agostini· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.