SignRefine is proposed, a sign language video generation model that produces comprehensible signing from 2D keypoint conditioning alone, generalizing across appearances and visual conditions, and is preferred by sign language users for visual quality and comprehensibility in more than 80% of comparisons.
Abstract
Sign language video generation demands precise hand and facial articulation, yet modern video diffusion models, trained predominantly on spoken-language video, produce artifacts that render signing unintelligible. We propose SignRefine, a sign language video generation model that produces comprehensible signing from 2D keypoint conditioning alone, generalizing across appearances and visual conditions. Our approach builds on a pretrained video diffusion transformer and introduces local adapters with spatial grounding to selectively refine hand and face regions, steering the strong base model's prior toward accurate articulation. To enable this work and support broader sign language research, we present NVSign, a large-scale dataset of video content natively produced in sign language, offering diverse signer appearances, environments, and natural conversational settings. Trained on this data, our model shows up to 30% improvement in hand pose precision metrics over the strongest baseline and is preferred by sign language users for visual quality and comprehensibility in more than 80% of comparisons.
SignMimic achieves state-of-the-art-level performance on video quality, identity similarity, and frame continuity while also achieving minimal loss when performing back translation (SLT) on generated videos.
Zhe-Wen He, Jun-Yi Yu, Hao Huang et al.· 0 citations
Sign language serves as the primary linguistic modality for the deaf and hard-of-hearing community. However, the vast majority of sign language content contained within historical video archives remains inaccessible — unindexed, unsearchable and effectively invisible to computational analysis. This paper proposes a nov...
Takashi Koyano· Journal of Digital Media Man...· 0 citations
A novel framework based on conditional Variational autoencoder for SLT (VSLT) that facilitates direct and sufficient cross-modal alignment between sign language videos and spoken language text is proposed, and a shared Attention Residual Gaussian Distribution (ARGD) which considers the textual information as a residual...
Rui Zhao, Liang Zhang, Biao Fu et al.· International Journal of Com...· 0 citations
Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unifi...
Zhaoyi An, Si-Han Tan, Youngbae Hwang et al.· 0 citations
SignSeek sets a new state-of-the-art performance in cross-corpus retrieval on ASL-Citizen, WLASL, and NMFs-CSL without any downstream fine-tuning, surpassing methods explicitly trained on BSL and outperforming prior skeleton-based methods.
Sobhan Asasi, Ozge Mercanoglu Sincan, Richard Bowden· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.