Skip to content
Preprint

ProtoAct: Turning Wet-Lab Protocols into Embodied Robotic Actions

Aug 2026 · 1 citation · 17 references
Computer Science

TL;DR

ProtoAct, a structured protocol-grounding framework that converts free-form biological procedures into state-aware, embodiment-ready action sequences, is presented, providing a practical interface between biological protocol understanding and embodied robotic execution.

Abstract

Biological wet-lab protocols are written for trained researchers and often leave routine operations, state-dependent conditions, and contextual parameters implicit, making them difficult to translate into robot-executable actions. We present ProtoAct, a structured protocol-grounding framework that converts free-form biological procedures into state-aware, embodiment-ready action sequences. ProtoAct uses ProtoRAG to retrieve manually annotated examples for context-sensitive parsing, employs RefineChecker to detect and revise missing or inconsistent steps, and applies ActSchema to map the refined procedure into constrained JSON function sequences. We further introduce BioP2E, for which we manually annotate 22 cell-culture protocols into 258 monitoring conditions, 910 executable subtasks, and 962 grounded action calls. Evaluation across seven large language models demonstrates that ProtoAct can be effectively instantiated with different backbones. Ablations confirm that retrieval, posterior checking, and schema constraints make complementary contributions. The parsed subtasks further support demonstration collection and VLA model training, enabling successful execution in both simulation and real-robot settings. ProtoAct thus provides a practical interface between biological protocol understanding and embodied robotic execution.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation

CodeActionBench is introduced, a benchmark of 25 manipulation tasks that evaluates this capability through agentic Code-as-Policy and provides a controlled testbed for measuring how general-purpose models translate their capabilities into manipulation behavior and for examining typical failure scenarios in that process...

Yiheng Lyu, Xueying Jiang, Wen-Hao Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Long-horizon robot manipulation reuses skills across many task compositions, but improving these compositions with additional end-to-end demonstrations is costly. A practical self-improving system must decide both what to teach next and where to apply that supervision. We present ROBOCOACH, a world-model-guided coachin...

Jia-Jun Liu, Yi-Fan Chen, Yi-Chao Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SAGE: Symbolic Action-Gating and Editing for LLM Task Planners

SAGE (Symbolic Action-Gating and Editing), a single-LLM planner built from two lightweight mechanisms: a domain-agnostic symbolic gate that blocks precondition-violating actions with typed reasons as a runtime safety monitor, and a local edit that regenerates only the failed sub-goal's suffix, keeping completed and unt...

T. Bui, Jongsul Moon, Youngouk Kim et al. · 0 citations
Preprint Sep 2026

In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks

We study robotic in-context learning (ICL), an emerging paradigm that enables robots to infer and execute tasks from visual demonstrations. Despite its growing promise, the problem itself remains under-defined: a visual demonstration simultaneously conveys action trajectories, object semantics, manipulation affordances...

Min-Xing Li, Ming-Hao Han, Wei-Zhi Zhao et al. · 0 citations
Nov 2026

DEPT-BT: Learning Parameterized and Revisable Behavior Trees From a Single Demonstration

Learning from Demonstration (LfD) combined with Behavior Trees (BTs) aims to lower the programming burden required to construct robot task programs from demonstrations. However, existing approaches have two structural limitations: skills are bound to specific object instances with no mechanism for runtime rebinding, an...

Bo-Gang Jiang, Peng-Ji Wu, Zhi-Jie Xu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CAPEX: Efficiently Distilling Foundation Model Behavior into Deployable Robot Policies through Experience-Adaptive Reasoning

CAPEX, an experience-conditioned demonstration collection framework that uses execution experience from previous attempts to adapt how frequently the foundation model must observe, reason, and replan, is introduced and suggests that foundation models can serve as scalable sources of reusable robot experience.

Shivam Aarya, Xi-Jia Zhang, Cheng-Yue Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.