Jul 2026· International Conference on Computer Vision, Al and Intelligent Automation· Vol 14260, pp. 142600O - 142600O-6· 0 citations· 10 references
Engineering
TL;DR
This paper systematically reviews the evolution of computer vision-based human motion recognition technology, from traditional handcrafted feature methods to convolutional neural networks and recurrent neural networks in the deep learning era, and then to the recently emerging Transformer architectures and vision-language models.
Abstract
Human motion recognition is an important research direction in computer vision, aiming to automatically analyze human pose changes and action semantics from images or videos captured by visual sensors. This paper systematically reviews the evolution of computer vision-based human motion recognition technology, from traditional handcrafted feature methods to convolutional neural networks and recurrent neural networks in the deep learning era, and then to the recently emerging Transformer architectures and vision-language models. The core ideas and technical characteristics of various approaches are comprehensively analyzed. On this basis, key challenges in current research are deeply explored, including the complexity of spatiotemporal feature extraction, robustness issues with occlusion and viewpoint changes, difficulties in fine-grained action and few-shot learning, and the trade-off between multimodal fusion and computational efficiency. Finally, future development trends are discussed, pointing out that cutting-edge directions such as 3D human pose estimation, fisheye lens adaptation, micro-action detection, and attribute-aware generation will drive the field toward higher precision, stronger generalization capabilities, and broader application scenarios. This paper aims to provide systematic technical reference and forward-looking insights for researchers in human motion analysis and intelligent human-computer interaction.
A compact and deployable Convolutional Neural Network-Long Short Term Memory (CNN-LSTM) framework that combines a 2D convolutional backbone (AlexNet) for frame-level descriptors with an LSTM head for sequence modeling is proposed, indicating a robust, real-time-capable solution for video understanding in both offline analytics and online deployment.
H. Khan, Altaf Hussain· ICCK Transactions on Advance...· 0 citations
A critical review of computer vision, illustrating how architectural design, learning paradigms, and evaluation practices have co-evolved over time to facilitate more flexible and scalable systems, and outlining new research directions.
Aiming at the key problems such as the separation of action recognition and quality assessment tasks, coarse feedback granularity and high labeling cost in the automatic evaluation of rehabilitation training, this paper proposes PDDS‑Net. The framework takes human skeleton sequence as input and realizes action classification and location-level deviation location synchronously through the collaborative architecture of a global action recognition stream and a local part evaluation stream. In the local flow, the joint nodes are divided into three functional part groups: upper limb, trunk and lower limb. independent graph convolutional subnetworks are used for decoupled coding, and a contrastive learning strategy is used to drive the attention map to focus on abnormal joints without frame‑level annotations. This is mainly achieved through the adaptive fusion mechanism to dynamically integrate the confidence of the deep network and the matching score of dynamic time warping template to maintain decision stability under conditions of pose degradation. Experiments on the NTU RGB+D and PKU‑MMD public datasets show that the action recognition accuracy of PDDS‑Net reaches 93.7%, the average precision of position‑level feedback is 0.656, 82.3% of the original AUC is still maintained at a 40% joint dropout rate, and the single‑frame inference latency is 41.8 ms, which meets the requirements of real-time interaction. This study provides a replicable technical path for the validation of rehabilitation evaluation algorithms without clinical data collection.
Mingxiang Yang· Journal of Discovery Core· 0 citations
The article compares the performance of traditional machine learning techniques with recent deep learning architectures such as CNNs, RNNs, TCNs, and Transformers, based on accuracy, computational cost, and suitability for real-world disorderly plotting.
Disha Deotale, Madhushi Verma, P. Suresh et al.· Discover Artificial Intellig...· 0 citations
A comparative analysis of existing studies is presented to highlight the evolution of deep learning techniques and their effectiveness in improving recognition accuracy and computational efficiency and emerging research directions are outlined to provide insights for future research.
Patel Bhautika Ronak· International journal of res...· 0 citations