2026· International Conference on Language Resources and Evaluation· pp. 9504-9513· 0 citations· 43 references
Computer Science
TL;DR
This work proposes a novel text-to-sign translation based on model pretraining, which enhances semantic alignment by inheriting codebook-oriented prior knowledge from masked self-supervised models.
Sign Language Production (SLP) plays a crucial role in bridging the communication gap between the Deaf community and broader society, functioning alongside Sign Language Translation (SLT) and Recognition (SLR). In addition to the limited scale of available data, research on Vietnamese Sign Language (VSL) is further hin...
D. Thanh, Thang Cap· International Conference on...· 0 citations
A novel framework based on conditional Variational autoencoder for SLT (VSLT) that facilitates direct and sufficient cross-modal alignment between sign language videos and spoken language text is proposed, and a shared Attention Residual Gaussian Distribution (ARGD) which considers the textual information as a residual...
Rui Zhao, Liang Zhang, Biao Fu et al.· International Journal of Com...· 0 citations
Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unifi...
Zhaoyi An, Si-Han Tan, Youngbae Hwang et al.· 0 citations
Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks. However, fine-tuning LLMs for Gloss-Free Sign Language Translation (GFSLT) remains a challenge. In this paper, we investigate how to effectively adapt LLMs to the GFSLT task. We show that there are two key issues that need to be...
Shi-Wei Gan, Xiao Liu, Ya-Feng Yin et al.· 1 citation
A scalable and modular SLP framework based on Sign-Pose-VQ-VAE model, designed for low-resource settings, achieves state-of-the-art performance among keypoint-based methods on the PHOENIX14T benchmark, attaining a BLEU-4 score of 10.03 and surpassing the previous best method by 0.67 points.
Suvajit Patra, Arkadip Maitra, Swami Punyeshwarananda et al.· Proceedings of the Thirty-Fi...· 0 citations
SeRV (Semantic-Aligned Residual Vector Quantization), a semantic-aligned RVQ tokenizer for ASL generation, achieves state-of-the-art pose accuracy on both How2Sign and YouTube-ASL datasets, while producing semantically consistent 3D ASL motion directly from text.
Hong-Yu Wu, Xu-Ying Wu, Tian-Hao Wu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.