Skip to content
Preprint

TANDEM: Task and Motion Planning with As-Needed Demonstrations for Efficient Vision-Language-Action Model Fine-tuning

Sep 2026 · 0 citations · 31 references
Computer Science

TL;DR

TANDEM (Tamp with As-Needed Demonstrations for Efficient Model fine-tuning), a system that combines TAMP with selective human teleoperation to collect demonstrations for tasks beyond the planner's capabilities, is presented.

Abstract

Human teleoperators spend substantial time demonstrating behaviors that robots can already perform autonomously, limiting the scalability of data collection for robot foundation models. Task and motion planning (TAMP) can automate many of these behaviors, but a fixed planning domain may not support every stage of a long-horizon manipulation task. We present TANDEM (Tamp with As-Needed Demonstrations for Efficient Model fine-tuning), a system that combines TAMP with selective human teleoperation to collect demonstrations for tasks beyond the planner's capabilities. Our key idea is to represent human assistance as an on-demand planning capability. Given a language instruction and visual observation, TANDEM uses pretrained vision-language models to extend the planning domain with missing predicates and human-executed magic operators. This allows the planner to interleave autonomous and human-executed stages without task-specific intervention points. After each human stage, TANDEM re-perceives the scene and checks whether the intended effects hold before resuming autonomous planning. To support fine-tuning vision-language-action (VLA) models, TANDEM also uses example pretraining trajectories to align planner-generated motions with the target model's pretraining distribution. We evaluate TANDEM on five long-horizon manipulation tasks beyond the TAMP domain's capabilities. On a representative long-horizon task, TANDEM collects 2.9x as many demonstrations as full-task teleoperation at the same human intervention time. Fine-tuning a pretrained \pi_{0.5}-DROID model on 20 TANDEM demonstrations per task increases average task success from 0% to 60% across the five tasks.

View source

Similar papers

Preprint Aug 2026

Evidence-Gated Task and Motion Planning with Vision-Language Models

Evidence Acquisition and Feasibility Gating (EAFG) is proposed, a framework that acquires visual evidence through VLM-generated exploratory subgoals and TAMP-based execution and applies a feasibility gate to decide whether to proceed with task planning, acquire further evidence, or halt.

Tsunehiko Tanaka, Matthew Stephenson, Alistair Macvicar et al. · 0 citations
Open access Aug 2026

Large Language Model-Driven Symbolic Planning for Long-Horizon Robotic Manipulation Tasks

VLA-SP (Vision-Language-Action via Symbolic Planning), a two-stage Embodied Vision-Language-Action framework, enabling fully automated robotic execution from speech and vision inputs is proposed, demonstrating the strong interpretability, executability, and cross-platform applicability of the framework.

Han-Zhuo Zhang, Jiahao Xu, Yi-Chen Xu et al. · 0 citations
#small language model Preprint Sep 2026

SkipVLA: Skipping VLA Steps with Classical Planning for Fast Robot Manipulation

SkipVLA is a hybrid policy that combines a pretrained VLA with a classical motion planner, using the planner for free-space motion and querying the VLA only for contact-rich skills such as grasping and placing, demonstrating up to 2.5x faster task completion and significantly lower energy consumption.

Kaivalya Agrawal, Md Ashiqur Rahman, Raymond A. Yeh et al. · 0 citations
#artificial intelligence Preprint Sep 2026

FRAMES: Failure Recovery And Monitoring of Embodied Skills for Humanoid Loco-Manipulation

Large language model (LLM) planners can decompose natural-language instructions and select reusable robot skills, but choosing the correct skill does not guarantee successful physical execution. This gap is especially important in humanoid loco-manipulation, where errors during approach, grasping, transport, or placeme...

Ajay Vikram Periasami, Xin Luo, Hao-Yu Li et al. · 0 citations
Aug 2026

MulPlanLM: multimodal robotic task planning with vision-language models and physical feedback

Experimental results in various task scenarios show that the proposed framework consistently improves overall task success rates compared with unimodal settings with different LLMs and achieves a higher success rate compared to using only visual or force data.

Young-Chae Son, Dong-Han Lee, Soo-Chul Lim · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.