Hybrid Spatial–Kinematic Learning for Robust Hand Gesture Recognition Under Transition Ambiguity
Abstract
Hand gesture recognition based on video data plays a key role in enabling natural and intuitive human–machine interaction. However, existing approaches often struggle with ambiguous gesture patterns, particularly during transitional states between actions, where visual similarity leads to frequent misclassification. This paper proposes a hybrid spatial–kinematic learning framework for robust hand gesture recognition under transition ambiguity. Unlike conventional approaches, the proposed method explicitly addresses ambiguity in intermediate gesture states by integrating spatial features from a lightweight CNN, temporal modeling using LSTM, and angle-based kinematic features derived from hand skeletons. Experimental results demonstrate that the proposed model achieves an accuracy of 94.2%, outperforming baseline CNN, ResNet, and CNN–LSTM models. In addition, the framework maintains real-time performance at 28 ms per frame, making it suitable for edge deployment. These results highlight the effectiveness of multimodal feature integration for improving robustness in gesture recognition under challenging real-world conditions. Furthermore, the lightweight design enables real-time inference, making the system suitable for deployment on edge devices in practical human–machine interaction applications.