A Robust Hybrid Gesture Recognition Framework for Real-Time Human-Robot Interaction
Abstract
This study presents a real-time vision-based gesture control framework for mobile robots using a monocular RGB camera. The goal is to enable intuitive and low-cost human– robot interaction without requiring specialized sensing hardware. The proposed system combines geometric rule-based reasoning with a Support Vector Machine (SVM) classifier through confidence-aware fusion. Scale-normalized hand-crafted features are extracted from 2D hand landmarks, and temporal filtering with finite-state control logic is introduced to suppress transient misclassifications and prevent unintended robot motion. The framework is evaluated on an expanded self-collected dataset containing 12,599 samples from five users under four environmental conditions. To avoid data leakage, quantitative evaluation is conducted using a Leave-One-User-Out protocol. Experimental results show that the hybrid framework achieves strong cross-user performance and outperforms lightweight baseline models. Sensitivity analysis further demonstrates the trade-off between confidence thresholding, temporal smoothing, and control responsiveness. Real-world deployment on a mobile robot confirms that the proposed framework can generate stable and responsive motion commands directly from hand gestures. Overall, this study demonstrates that reliable gesture-based robot control can be achieved using lightweight vision algorithms, providing a practical solution for accessible human–robot interaction.