Aug 2026· International Conference on Digital Image Processing· Vol 14351, pp. 143512V - 143512V-9· 0 citations· 21 references
Engineering
TL;DR
A vision-guided robotic action generation framework that explicitly models the data flow from visual perception to robotic action execution and introduces a structured visual data extraction mechanism that interprets raw visual outputs into type-consistent, constraint-aware, and physically feasible motion parameters, enabling reliable perception-action coupling in industrial assembly systems.
Abstract
Visual perception plays a critical role in industrial assembly systems, where robotic actions are driven by image-based sensing under variable and data-dependent conditions. A key challenge in such systems lies in transforming unstructured visual inference outputs, including object detection and pose estimation results, into structured and executable control parameters that can be reliably grounded in physical execution. In practice, mismatches between perception outputs and downstream control interfaces often lead to execution errors and reduced system robustness in multi-stage assembly processes. To address this challenge, this paper proposes a vision-guided robotic action generation framework that explicitly models the data flow from visual perception to robotic action execution. The framework introduces a structured visual data extraction mechanism that interprets raw visual outputs into type-consistent, constraint-aware, and physically feasible motion parameters, enabling reliable perception-action coupling in industrial assembly systems. By decoupling visual interpretation from low-level control execution, the proposed approach improves modularity and robustness across heterogeneous hardware platforms and low-code industrial orchestration environments. The proposed framework is implemented and evaluated through an end-to-end, data-dependent toy vehicle assembly task involving multiple perception-driven operations. Experimental results demonstrate that the proposed method significantly improves perception-action alignment robustness, achieving higher phase-level execution reliability and an end-to-end assembly success rate of up to 94%, outperforming baseline approaches that lack explicit visual data alignment mechanisms.
A DNN-based vision-guided robotic assembly framework that integrates computer vision, intelligent decision-making, and real-time robotic control is proposed, supporting flexible automation and next-generation smart manufacturing in Industry 4.0 environments.
O. Dahl, K. Nygaard· International Journal of Int...· 0 citations
The outcomes demonstrate the efficacy of combining edge intelligence with closed-loop robotic control by confirming consistent behavior throughout simulation and limited physical testing.
Xiaoming Liu, Wei Su, Jie Zhang et al.· Scientific Reports· 0 citations
Experimental results in various task scenarios show that the proposed framework consistently improves overall task success rates compared with unimodal settings with different LLMs and achieves a higher success rate compared to using only visual or force data.
Robots that follow open-ended language instructions need to connect semantic intent to visual scene understanding, geometric feasibility, object states, and physical interaction conditions. End-to-end Vision-Language-Action policies have improved cross-task generalization, but they typically map visual and language inputs directly to robot actions, leaving limited explicit structure for long-horizon decomposition, physical verification, and recovery. We present \method, a zero-shot hierarchical framework functionally inspired by the division of roles in the human brain, comprising visual perception and state inference, language grounding and action-sequence generation from a shared atomic action library, cost-based plan selection, and real-robot execution and verification. The framework grounds commands in explicit object states, composes reusable atomic actions into task-conditioned sequences, ranks alternative sequences by execution cost, and verifies intermediate physical outcomes from refreshed observations. In the evaluation, \method{} completes 10/10 clean board trials, 10/10 pick-and-place trials, and 4/5 pyramid stacking trials for both the flat and irregular initial-layout conditions; the corresponding mean task progress is $99.03\%$, $100.00\%$, and $96.67\%$ respectively. Across all evaluated conditions, \method{} achieves higher success rates than ReKep, Dream2Flow, and $\pi_{0.5}$ benchmarks, demonstrating the effectiveness of combining explicit object-state reasoning, compositional atomic actions, cost-based plan selection, and closed-loop execution verification.
IMBENCH is introduced, a benchmark designed to evaluate intuitive manipulation as an integrated capability spanning perception, physical reasoning, action generation, and iterative execution, and position IMBENCH as a step toward evaluating and enabling more integrated, adaptive physical intelligence.
A shared-autonomy framework that assists the operator throughout this process ofTeleoperating a robotic manipulator in industrial environments demands precision that camera-based interfaces alone struggle to deliver and is validated on a quadruped mobile manipulator.
Murilo Vinicius da Silva, Ricardo V. Godoy, Juliano Negri et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.