Aug 2026· International Journal of Electronics and Communication Engineering· Vol 13, pp. 127-139· 0 citations· 26 references
TL;DR
The results indicate that introducing motion-consistent segmentation and structured decision fusion seems to be a good way for updating the CSLR systems beyond simply endwise paradigms.
Abstract
Continuous Sign Language Recognition (CSLR) has always stood quite difficult due to issues like coarticulation effects, availability of weak temporal annotations, and large variability among different signers, especially in a limited-resource scenario such as Indian Sign Language (ISL). Most of the current methods use end-to-end sequence modeling, which leaves out the explicit temporal structure and does not work well with weak supervision. Here, we propose a structured CSLR system that combines motion-guided segmentation, multimodal representation learning, and position-aware decoding aimed at overcoming these issues. More importantly, we present an MGPT-based temporal segmentation method that uses optical-flow-driven motion signals and Gaussian peak modeling to separate continuous signing sequences into consistent motion segments, which results in the reduction of transitional ambiguity. The spatial-temporal features are obtained with the help of a dual-stream architecture that integrates ResNet50-based visual representations and skeleton keypoint features, being then temporally modeled by a multi-layer LSTM network. To improve sequence-level consistency, we also introduce a Word Position Graph (WPG) for structured decoding along with Gaussian-weighted frame voting to highlight informative temporal regions and, at the same time, downplay noisy transitions. The approach we suggested was tested on the ISL-CSLRT dataset with weak sentence-level supervision. The experimental results show that our framework reaches 92% accuracy and a Word Error Rate (WER) of 0.07, greatly beating the baseline voting strategies. Statistical verifications, including multi-run evaluation and significance testing, have confirmed the robustness of the improvements. Also, comparing with representative CSLR methods has shown that the method of explicit temporal segmentation and position-aware decoding is very effective, especially when the dataset is scarce. Besides, the results indicate that introducing motion-consistent segmentation and structured decision fusion seems to be a good way for updating the CSLR systems beyond simply endwise paradigms.
This work repurposes the CTC decoder to restrict representation-based scoring to decoder-aligned gloss regions, rather than exposing the acquisition function to the entire unfiltered video, and introduces RAIDAL, which achieves its strongest data-efficiency gains over competing baselines in large-vocabulary, budget-limited settings, while remaining competitive in the smaller-vocabulary, large-budget setting.
R. A. Diniz Augusto, Gabriel L. Oliveira, Erickson R. Nascimento· 0 citations
With the growing emphasis on accessibility-oriented technologies and inclusive intelligent systems, Continuous Sign Language Recognition (CSLR) has attracted increasing attention as a key technique for bridging communication between Deaf and hearing communities. However, existing methods still suffer from insufficient exploitation of visual information, weak temporal alignment supervision, and inconsistency between training and inference, making it difficult to jointly improve recognition accuracy and model robustness. To address these issues, we propose TPA-Seq2Seq (Tri-Prior Aligned Seq2Seq), a tri-prior enhanced Seq2Seq framework for character-level CSLR. Specifically, we introduce PGF (Pose-Guided Fusion) to extract hand-arm-face keypoint descriptors offline through the MediaPipe pipeline and fuse them with RGB-based temporal semantics in a lightweight manner, thereby supplementing fine-grained visual priors. We further design ATAL (Auxiliary Temporal Alignment Loss), a CTC-based auxiliary constraint imposed on the encoder side to strengthen explicit temporal alignment supervision. In addition, we propose SSF (Scheduled Semantic Forcing), a piecewise teacher-forcing decay strategy that alleviates exposure bias and improves generalization. Experiments on the CSL dataset show that TPA-Seq2Seq reduces WER compared with the ResNet18-LSTM baseline and maintains consistent performance across different random seeds.
Ya-Han Yang, Rui Wang, Xiao-Fang Li et al.· International Conferences on...· 0 citations
Self-supervised sign language representation learning must model two properties not central to natural-image SSL: signs are produced by a small set of anatomically distinct articulators, and their meaning depends on the temporal organisation of those articulators. We introduce SignDino, a self-supervised sign-video encoder that moves the DINOv3 student--teacher recipe from the spatial domain of image crops to the temporal domain of tracked sign streams. Each video is decomposed into left-hand, right-hand, and face streams by a detector-first YOLOv8n+ByteTrack pipeline. A frozen DINOv3 ViT-B/16 embeds each per-frame anatomical crop, while lightweight temporal Transformers, not the image backbone, form the student and EMA teacher. They are trained by temporal DINO self-distillation, frame-level masked-token prediction in the style of iBOT, KoLeo feature spreading, and Gram anchoring of the frame-to-frame similarity structure. This design keeps strong image-level visual primitives fixed and learns only how articulator states evolve across time. We evaluate on sign-to-English translation, isolated sign recognition, and fingerspelling detection benchmarks. Across these tasks, SignDino provides a strong public self-supervised representation and shows competitive or state-of-the-art performance under matched downstream evaluation.
Jun-Yi Hu, Zhe-Wen He, Hao Huang et al.· 0 citations
Sign language is an essential communication system for hearing-impaired individuals, which mainly depends upon complex hand gestures and facial expressions. Automating Sign Language Recognition (SLR) from videos can enhance barrier-free communication, yet it remains challenging due to the subtle nature of signs in diverse environments. Existing modules often need extensive manual feature extraction and struggle with real-time applications due to high latency. Thus, an effective model for SLR using videos is presented, named Supercell Thunderstorm Paper Publishing-based Optimization enabled Convolutional Grid Long Short-Term Memory (STPPO_CGLSTM). The frames are extracted from the input video. Then, the Arithmetic Mean Filter is employed for pre-processing the extracted frames. Moreover, humans are segmented by employing correlational spectral clustering. Thereafter, hand action unit detection is carried out by considering the AU-Net. Following this detection task, effectual features are extracted. Lastly, sign language is recognized by utilizing CGLSTM, where the hyperparameters of CGLSTM are trained using STPPO. The analytical measures, namely, accuracy, Positive Predictive Value (PPV), Negative Predictive Value (NPV), and False Omission Rate (FOR), obtained better outcomes for STPPO_CGLSTM, which is 96.707%, 98.219%, 94.016% and 5.984% using k-fold cross validation.
A. Babisha, G. Srikanth· Discover Computing· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.