Skip to content
Conference

LVR-Draw: A Language-Vision Pipeline for Robotic Drawing with Interactive Scene Verification and Correction

Aug 2026 · 2026 IEEE International Conference on Mechatronics and Automation (ICMA) · pp. 1141-1146 · 0 citations · 22 references

Abstract

This paper presents LVR-Draw, a fully local language–vision pipeline for robotic drawing that integrates structured scene generation, multimodal verification and correction, and deterministic execution within a unified Human–AI–Robot loop. Given a natural language prompt, a Large Language Model (LLM) generates a structured scene representation in a predefined format, which is rendered into an interpretable image. A Vision–Language Model (VLM) then performs visual inspection to detect inconsistencies in object placement and spatial relationships. These observations are processed by the LLM to produce structured editing operations, enabling iterative refinement of the scene. After validation, the refined scene is converted into executable robot instructions through a deterministic pipeline, supporting predictable and reproducible execution without using generative models for control. Experimental results suggest that the system can generate valid scene representations, support multimodal correction, and preserve drawing order during physical execution.

View source

Similar papers

Open access Jul 2026

NEO: NeRF It Once, Edit It Many Times for Continuous Object Manipulation

In this letter, we present NEO, a unified framework providing language-guided NeRF editing for robotic manipulation. Our letter introduces (i) a language-guided object removal that combines neural field resampling with multiview-consistent progressive inpainting, (ii) a direct NeRF weight editing method utilizing knowledge distillation, composing original and edited NeRFs via a teacher–student model, enabling coherent modeling of future scene states before a robot executes an action, and (iii) the first benchmark (NEO-Dataset) for quantitatively evaluating NeRF scene editing methods suitable for robot manipulation. We show that our approach outperforms state-of-the-art baselines in scene editing tasks, including object removal and pick-and-place robotic experiments, yielding visually coherent and geometrically consistent edits that reduce artifacts commonly introduced by prior methods. Finally, we showcase the capability of NEO for multi-stage robotic assembly tasks by preserving a persistent NeRF plus language-field representation after each edit, enabling iterative future-state scene representation prediction without requiring additional scene re-scanning.

M. Zieliński, David Hall, Dominik Belter et al. · 0 citations
Preprint Sep 2026

IM-ENGINE: Image Editing for Embodied Data Generation

Learning-based manipulation requires supervision that is both semantically meaningful and physically executable, but current data pipelines often provide only one of these properties. Human demonstrations capture intent but are costly to collect and constrained by the human-robot embodiment gap, while simulation can scale data generation but often under-specifies functional behavior. We present IM-ENGINE, a simulator-grounded pipeline that uses image editing as an intermediate representation for embodied data generation. Given a rendered scene with known geometry, depth, segmentation, and camera parameters, IM-ENGINE edits the image to inject task-relevant semantics, recovers explicit 3D state using simulator priors and an unchanged anchor object, refines the state in physics, and converts it into robot-executable supervision. We instantiate the pipeline for dexterous grasp synthesis and goal-state generation. For grasping, IM-ENGINE generates a human grasp in image space, recovers the hand-object interaction, retargets it to a robot hand, and refines it into physically validated robot grasps. For goal generation, it edits a rendered scene into a desired outcome, recovers the target-object pose, and refines it into physically valid, semantically meaningful goals and trajectories. This combination of generative semantic priors and simulator grounding enables scalable task-relevant supervision for robot learning.

Yian Wang, Jun-Yi Cao, Xiao-Wen Qiu et al. · 0 citations
Book Open access Jul 2026

Natural-Language to Geometry Diagrams: A Constraint-Based Pipeline for Precise Visual Reasoning

GenGX is a system that generates precise geometric diagrams from natural-language descriptions by combining large language model (LLM) interpretation with symbolic constraint solving by combining large language model (LLM) interpretation with symbolic constraint solving.

Kavi Wilson, P. Todd · 0 citations
Conference Aug 2026

Scene-Level Planning for Temporally Coherent AI Video Generation Using Multimodal Representations

Recent text-to-video systems can generate visually appealing clips from natural language prompts, yet narrative prompts often contain multiple implicit temporal stages that require the generator to infer scene decomposition, subject persistence, action ordering, and visual continuity from a single unstructured input. This frequently leads to temporally inconsistent or structurally ambiguous outputs. In this work, we investigate whether introducing an explicit scene-planning layer can improve multi-stage video generation. We compare three generation paradigms: direct single-prompt generation, naive prompt decomposition, and a structured scene-planning pipeline. The proposed approach first converts a narrative prompt into a lightweight structured scene representation containing global subject information, visual style constraints, and scene-level descriptions, which is then compiled into scene-conditioned prompts for sequential video generation. A reference-guided continuation mechanism conditions the second scene on the final frame of the first scene to improve cross-scene identity and visual continuity. The framework is generator-agnostic and can operate on modern text-to-video backends without modifying the underlying models. To evaluate generation quality, we adopt a vision-language model (VLM) as an automatic judge that assesses prompt relevance, temporal continuity, aesthetic quality, and narrative clarity across candidate videos. This study provides an empirical investigation of how structured intermediate planning influences narrative video generation and offers a lightweight framework for improving temporal coherence in AI-generated videos.

Jing Chen · 0 citations
Preprint Aug 2026

CoT-Edit: Let CoT Guide Instruction Video Editing

Text-driven instruction-based video editing in complex scenes remains challenging: purely textual prompts often fail to capture precise spatial relationships and physical constraints, resulting in target ambiguity and physically implausible outcomes. To address this, we propose a plan--guide--edit framework that explicitly bridges semantic intent and spatial execution. In our framework, a Chain-of-Thought (CoT)-enhanced multimodal large language model (MLLM) serves as a planner, performing structured reasoning over the video and instructions to derive a precise sequence of bounding boxes and attribute-enriched editing directives. These spatial priors then guide a box-conditioned mask generator, transforming ambiguous global retrieval into localized, context-aware refinement and producing masks that more accurately capture object scale, contact relationships, and placement. Building on these spatial and semantic signals, a diffusion-based editor integrates the masks, enriched instructions, and frame features to render high-fidelity edits that remain temporally coherent and spatially well aligned. Trained first in a modular manner and then jointly, our framework achieves superior performance with reduced data requirements, delivering precise localization in scenes with multiple similar objects and physically consistent object additions, and extensive experiments demonstrate state-of-the-art performance over multiple strong baseline methods. More details are available at: https://github.com/flying-sky999/CoT-Edit

Sen Liang, Fengbin Guan, Youliang Zhang et al. · 5 citations
Preprint Aug 2026

Beyond Placement and Articulation: Usage-Driven Code Scenes for Embodied Interaction

RoomWright is presented, an agentic usage-driven framework for generating 3D scenes represented entirely as code for embodied interaction, providing interactive environments for embodied AI and policy learning.

Zijian Xiao, Zipeng Ye, Jin-Kun Hao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.