Skip to content
Preprint

SSC: A Verifiable Structured Representation for Bimanual Manipulation Labelling

Aug 2026 · 1 citation · ⚡ 1 influential · 28 references
Computer Science

TL;DR

The Structured Subtask Chain (SSC), a state-transition representation that bridges these extremes, is proposed and instantiate the pipeline on BEHAVIOR-1K for logic verification and content completion.

Abstract

Subtask labels decompose a long-horizon manipulation demonstration into shorter semantic segments for policy training and evaluation. Natural language descriptions are easy to read, but their linguistic variability makes automatic verification difficult. Rigid template formats, such as BEHAVIOR-1K's skill_annotation, are linguistically over-segmented, hindering both readability and annotation consistency. We propose the Structured Subtask Chain (SSC), a state-transition representation that bridges these extremes. A demonstration is a sequence of Structured Subtask Template (SST) entries. Each SST stores core action components (subject, predicate, object), flexible conditions (adverbial modifiers such as spatial or instrumental phrases), a base-motion field separate from arm actions, and an after-state scene graph. Built on this format, SSC supports three vision-language assisted functions: rendering SSTs as natural language, checking the assembled chain against four state-transition rules, and completing underspecified fields through a query resolution cascade. We instantiate the pipeline on BEHAVIOR-1K (50 tasks, 3 episodes per task, 2,357 annotated action cells) for logic verification and content completion, evaluating 13 selected state-of-the-art VL models as candidate verifiers and reporting labelling anomalies.

View source

Similar papers

Preprint Sep 2026

X-Planner: Event-Structured Task Planning for Embodied Intelligence

This work presents X-Planner, a planning front-end that addresses both the supervision and representation of embodied reasoning, and describes planning-text quality and downstream execution.

Howard Lu, Shalfun Li, Porter Pan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SAGE: Symbolic Action-Gating and Editing for LLM Task Planners

SAGE (Symbolic Action-Gating and Editing), a single-LLM planner built from two lightweight mechanisms: a domain-agnostic symbolic gate that blocks precondition-violating actions with typed reasons as a runtime safety monitor, and a local edit that regenerates only the failed sub-goal's suffix, keeping completed and unt...

T. Bui, Jongsul Moon, Youngouk Kim et al. · 0 citations
Preprint Aug 2026

BIMScript: Material-Aware Structured Scene Programs for BIM Ingestion

Structured-language models such as SceneScript reconstruct a scene as a short program of parametric commands, an inherently editable and semantically explicit representation. We ask three questions that stand between such models and their most compelling application, automated ingestion of existing buildings into BIM t...

P. Naikade, Thomas B. Moeslund, Andreas Møgelmose · 0 citations
Preprint Aug 2026

PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans

Civil infrastructure compliance checking has long relied on engineers manually reading legacy 2D plans; however, OCR-based automation strips away the geometry and layout essential for interpreting these plans. We present a Visual-First Multimodal Retrieval-Augmented Generation (RAG) framework called PlanSightRAG. It in...

N. Subedi, S. Datta, A. Abdelaty et al. · 0 citations
Open access 2026

STRUDEL: Unrolling a Benchmark for Evaluating Vision-Language Models on Structured Diagram Understanding across Domains

STRUDEL establishes a scalable foundation for assessing and advancing VLMs torward deeper and more systematic understanding of structured visual information across domains, and reveals that models excel at association tasks, yet struggle with quantification and identification, where precise structural understanding is...

Daniel Steinigen, Lucie Flek, Sebastian Houben · 0 citations
Preprint Sep 2026

VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement

VeriPhy, an auditable physical-verification system in which a text-only planner compiles the prompt into typed physical obligations and a statically validated execution plan before any frame is observed, is presented.

Wenzhuo Xu, Yu-Chen Zhu, Chongjian Ge et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.