Skip to content

Evaluation of transfer learning by means of fuzzy logic for hand gesture recognition

Jul 2026 · Multimedia tools and applications · Vol 85 · 0 citations · 42 references
Computer Science

TL;DR

This work introduces a systematic evaluation framework for static HGR that combines transfer learning with a fuzzy logic expert system to provide a reproducible and transparent ranking of candidate models, and an evidence-based comparison of classical versus modern CNN architectures in static HGR.

View source

Similar papers

Conference Jul 2026

Hybrid Spatial–Kinematic Learning for Robust Hand Gesture Recognition Under Transition Ambiguity

Hand gesture recognition based on video data plays a key role in enabling natural and intuitive human–machine interaction. However, existing approaches often struggle with ambiguous gesture patterns, particularly during transitional states between actions, where visual similarity leads to frequent misclassification. This paper proposes a hybrid spatial–kinematic learning framework for robust hand gesture recognition under transition ambiguity. Unlike conventional approaches, the proposed method explicitly addresses ambiguity in intermediate gesture states by integrating spatial features from a lightweight CNN, temporal modeling using LSTM, and angle-based kinematic features derived from hand skeletons. Experimental results demonstrate that the proposed model achieves an accuracy of 94.2%, outperforming baseline CNN, ResNet, and CNN–LSTM models. In addition, the framework maintains real-time performance at 28 ms per frame, making it suitable for edge deployment. These results highlight the effectiveness of multimodal feature integration for improving robustness in gesture recognition under challenging real-world conditions. Furthermore, the lightweight design enables real-time inference, making the system suitable for deployment on edge devices in practical human–machine interaction applications.

Vo Tu Duong, An Minh Luong, Ta Nguyen Duc Dung et al. · 0 citations
Open access Aug 2026

A Lightweight Dynamic Gesture Recognition Model Driven by Meta-Learning Under Small Sample Conditions

Dynamic gesture recognition under small-sample conditions remains challenging due to the large variations caused by users, viewpoints, motion patterns, and hand configurations. This study proposes a lightweight dynamic gesture recognition framework that integrates meta-learning, Neural Architecture Search (NAS), and knowledge distillation (KD) to achieve rapid adaptation with limited training samples. Different from conventional recognition methods that mainly focus on feature extraction and model optimization, this study further considers the intrinsic symmetry characteristics of hand gestures. A mirror-aware skeleton representation strategy is introduced by modeling the structural correspondence between left- and right-hand keypoints, which reduces the distribution differences caused by hand-side variations and improves the generalization ability under few-shot conditions. The proposed framework adopts a lightweight spatio-temporal feature extraction module, an optimization-based meta-learning strategy, and a knowledge distillation mechanism to balance recognition accuracy and computational efficiency. Experiments are conducted on DHG-14, SHREC2017, FPHA, LMDHG, and 20BN-Jester datasets under different few-shot settings, including cross-user variation, viewpoint variation, motion speed variation, partial occlusion, and background interference. The experimental results demonstrate that the proposed method achieves competitive recognition performance while significantly reducing model complexity and computational cost. The proposed framework provides an efficient solution for dynamic gesture recognition in resource-constrained scenarios.

Yaxu Xue, Feifei Ru, Jiawu He et al. · 0 citations
Open access Aug 2026

Hand Gesture Recognition Based on Multi-Scale Attention Graph Convolutional Network

Advances in artificial intelligence have made hand gesture recognition an important human–computer interaction modality. Graph convolutional networks (GCNs) are widely used for skeleton-based hand gesture recognition, yet their performance can be limited by weak semantic topology modeling, underused feature channels, and shallow spatio-temporal fusion. We propose a Multi-scale Attention Graph Convolutional Network (MA-GCN) that combines three components within one skeleton framework: a hybrid topology that augments physiological connections with semantic priors; a Gaussian Multi-Scale Channel Attention (GMCA) module for coordinate denoising and adaptive channel weighting; and a Local-Global Fusion Module (LGFM) that combines local convolutional features with channel-wise global attention. Ablation studies quantify the independent and joint contributions of these components. MA-GCN obtains Top-1 accuracies of 97.50%/95.95% on SHREC’17 Track and 94.29%/92.86% on DHG14/28 for the 14-/28-class settings. In a SHREC’17 Track-to-FPHA pre-train-then-fine-tune evaluation, it reaches 94.09% Top-1 accuracy, providing preliminary evidence that the proposed framework maintains effectiveness under cross-dataset transfer.

Xiaowei Han, Tingshan Yan, Yunjing Lu et al. · 0 citations
Review Open access 2026

Real-Time Hand Gesture Detection with 3-D Key Points Using Deep Learning

The importance of touchless systems for easier Human Machine Interaction (HMI) has been brought to light by recent pandemics. A key component of HMI, gesture detection represents a distinct class of useful computer vision applications. Outstanding outcomes in video processing are demonstrated by Convolutional Neural Networks (CNN) for a range of computer vision applications. This study reviews and compares various CNN versions for gesture detection along with their uses. To determine the correctness of the model, a local dataset of hand gestures is built and tested in conjunction with a synthetic gesture dataset. Using multilayered CNN, palm detection in a video frame is accomplished with efficient foreground-background separation. In a parametric research, these Mean Square Error (MSE) values are then contrasted with the known comparable system. The MSE of 12.6% obtained from the combination of synthetic and local datasets is better than that of recent research. In order to achieve 3 Dimensional (3D) estimation of a hand gesture with the aid of palm recognition, the number of key points allocated for the skeletal representation of the hand is crucial. This further enables the application of the technique in numerous upcoming genres.

Sanjay R Pawar, Rameez Shamalik, Nazim Mahammad Shaikh et al. · 0 citations
Jul 2026

Support Vector Machine Hyperparameter Optimization for Speech Command Classification Using Mfcc Features

This study proposes speech command classification using MFCC features and SVM with GridSearchCV hyperparameter optimization. Evaluating RBF/linear kernels, C (0.1-100), and gamma (0.001-scale) on Google Speech Commands Dataset (8 classes), the optimal configuration (RBF, $\mathbf{C}=\mathbf{1 0}$, gamma=0.01) as the best configuration, achieving a mean cross-validation accuracy of 83.15%. Evaluation on the independent test set yielded a final classification accuracy of 83.05%, with per-class F1-scores ranging from 0.73 to 0.89 (stop/up/yes). While lower than 3D CNN approaches (89.16%), the optimized SVM offers superior computational efficiency and provides a computationally efficient alternative compared to deep learning approaches with rigorous hyperparameter tuning as a practical baseline for lightweight speech command recognition.

Santoso, T. Sardjono, D. Purwanto · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.