Jul 2026· International Conference on Control, Decision and Information Technologies· pp. 1853-1858· 0 citations· 19 references
Abstract
Over the last decade, significant advances have been made in the field of human robot interaction (HRI), such as collaborative workcells in smart factories and autonomous mobile robots. However, for these systems to achieve widespread societal acceptance, more natural and efficient communication interfaces must be developed. This work proposes a computer vision-based human–robot cooperative system, composed of a set of gestures that, when combined into composite sequences, allow intuitive interaction with the environment. Gesture identification relies on a human pose analysis model based on deep learning, which earned the 1st place in the Flying Robots Demo at RoboCup 2025. The semantic and contextual interpretation of these sequences is conducted by Large Language Models (LLMs). Integration between visual perception and linguistic understanding enables a more expressive and adaptable form of interaction. The developed system achieved a high-fidelity gesture detection model, with average precision and recall of 0.94.
This study presents a real-time vision-based gesture control framework for mobile robots using a monocular RGB camera. The goal is to enable intuitive and low-cost human– robot interaction without requiring specialized sensing hardware. The proposed system combines geometric rule-based reasoning with a Support Vector Machine (SVM) classifier through confidence-aware fusion. Scale-normalized hand-crafted features are extracted from 2D hand landmarks, and temporal filtering with finite-state control logic is introduced to suppress transient misclassifications and prevent unintended robot motion. The framework is evaluated on an expanded self-collected dataset containing 12,599 samples from five users under four environmental conditions. To avoid data leakage, quantitative evaluation is conducted using a Leave-One-User-Out protocol. Experimental results show that the hybrid framework achieves strong cross-user performance and outperforms lightweight baseline models. Sensitivity analysis further demonstrates the trade-off between confidence thresholding, temporal smoothing, and control responsiveness. Real-world deployment on a mobile robot confirms that the proposed framework can generate stable and responsive motion commands directly from hand gestures. Overall, this study demonstrates that reliable gesture-based robot control can be achieved using lightweight vision algorithms, providing a practical solution for accessible human–robot interaction.
Chuyu Guo· 2026 IEEE International Conf...· 0 citations
Collaborative robot or COBOT has emerged as a key advancement in robotics, significantly transforming industrial automation. Consequently, real-time gesture recognition for human-COBOT interaction has recently garnered significant attention. However, both datasets designed for practical robot control scenarios and end-to-end deployment frameworks with seamless robot communication remain underexplored. To address these gaps, this paper makes two main contributions. First, we introduce CoboGestureV2, a novel dataset of 2362 gesture samples across 16 classes, specifically designed for continuous gesturebased COBOT control in practical workflows, covering both static and dynamic gestures including movement and rotation commands. Second, we present an end-to-end real-time gesture recognition framework deployed on an edge device (Jetson AGX Orin) and integrated with an MQTT-based communication protocol that enables lightweight, scalable message exchange between the recognition module, a coordination block, and the robot controller. On the proposed dataset, the framework achieves a frame-wise accuracy of 92.60%, edit score of 88.85% and TAL score of $\mathbf{7 1. 8 2} \boldsymbol{\%}$ in an offline setting, and $\mathbf{8 3. 6 5 \%}, \mathbf{7 7. 9 4 \%}$ and 48.93% for the respective metrics with a recognition latency of 0.87 seconds in a real-time online deployment.
Ha-Anh Nguyen, A. Nguyen, Duy-Anh Doan et al.· IEEE International Conferenc...· 0 citations
Robust human-robot interaction in complex environments requires accurate gesture perception, semantic scene understanding, and reliable task planning under limited onboard computing resources. This paper presents a cloud-edge multimodal interaction framework that integrates an enhanced YOLO-based gesture detector with coordinated large language model (LLM) and vision-language model (VLM) agents. The proposed detector, incorporates the Convolutional Block Attention Module (CBAM) into the neck and replaces the baseline bounding-box regression objective with Distance-IoU (DIoU) loss. These modifications improve feature discrimination and localization for small or partially occluded gestures in complex backgrounds. The cloud layer performs gesture detection, scene understanding, multimodal fusion, and action planning, whereas the TonyPi robot locally handles data acquisition, communication, action execution, and feedback. Experiments on a public gesture dataset and a custom dataset show that YOLO-DC achieves precision values of 98.9% and 95.0%, with mAP@0.5 values of 90.7% and 92.7%, respectively. System-level evaluation yields success rates of 95%, 88%, and 82% for single-action, composite-action, and vision-dependent tasks. A 30 participant evaluation yields an overall mean satisfaction score of 3.69 out of 5. These results demonstrate the feasibility of combining refined gesture detection with multimodal agents for resource-constrained robotic interaction.
Enabling humanoid robots to respond to human speech with synchronized and semantically meaningful gestures is fundamental to natural human-robot interaction. However, this task faces three critical barriers: the scarcity of semantically rich datasets, the"modality eclipse"where models ignore audio cues in favor of kinematic inertia, and the sim-to-real gap regarding physical safety. We propose RoboGesture, a robot-centric framework that co-designs data, modeling, and control to power a complete interactive human-humanoid system in which the robot listens, responds, and gestures in real time. We first establish the RoboGesture dataset featuring over 300 gesture categories and develop an automated pipeline to synthesize large-scale collision-free, robot-specific audio-motion pairs. Our architecture features a Hierarchical Semantic-Acoustic Aligner that extracts multi-granular prosodic and semantic cues directly from raw audio tokens. These cues drive a Streaming Conditional Motion Generator based on a diffusion transformer with conditional flow matching. To ensure high responsiveness, we introduce Anti-Inertia CFG Masking, which prevents the model from collapsing into repetitive historical patterns by compelling it to proactively mine control signals from the audio modality. Finally, an MPC-based safety filter ensures real-time, collision-free execution on physical hardware. Experiments on a Unitree G1 humanoid demonstrate that RoboGesture generates safer, more rhythmic, and more semantically appropriate responses compared to state-of-the-art baselines.
Zi-Fan Wang, Ziang Ren, Pengteng Shi et al.· 0 citations
Vision Language Models (VLMs) enable robots to visually perceive their environment as well as the actions and characteristics of their conversation partner or humans in collaboration. Especially for social robots deployed in everyday settings and for uncomplicated, natural use, it is essential that the robot has an understanding of situations that is appropriate to human customs. This paper presents initial experiences with the application of a Mistral AI language model with a Pepper robot for Human-Robot Interaction (HRI) in dialogue, as well as an investigation of the effects of additional visual information on response time in different models. The results show that incorporating visual information adds context to the dialogue with only a moderate increase in response time, enabling both the robot and the human to take into account unspoken elements of the situation. Furthermore, using an LLM hosted in Europe offers a solution that complies with European data protection regulations and can therefore facilitate real-life applications more easily.
Preliminary experiments on several representative arm gestures indicate that the proposed method can produce meaningful imitative motions from monocular RGB input only, while also highlighting limitations in more complex poses and wrist-related movements.
Anastasiya Ihnatovich, Igor Farkas· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.