Skip to content
Preprint

SignRefine: Adapting Foundational Video Models for Sign Language Generation

Sep 2026 · 0 citations · 82 references
Computer Science

TL;DR

SignRefine is proposed, a sign language video generation model that produces comprehensible signing from 2D keypoint conditioning alone, generalizing across appearances and visual conditions, and is preferred by sign language users for visual quality and comprehensibility in more than 80% of comparisons.

Abstract

Sign language video generation demands precise hand and facial articulation, yet modern video diffusion models, trained predominantly on spoken-language video, produce artifacts that render signing unintelligible. We propose SignRefine, a sign language video generation model that produces comprehensible signing from 2D keypoint conditioning alone, generalizing across appearances and visual conditions. Our approach builds on a pretrained video diffusion transformer and introduces local adapters with spatial grounding to selectively refine hand and face regions, steering the strong base model's prior toward accurate articulation. To enable this work and support broader sign language research, we present NVSign, a large-scale dataset of video content natively produced in sign language, offering diverse signer appearances, environments, and natural conversational settings. Trained on this data, our model shows up to 30% improvement in hand pose precision metrics over the strongest baseline and is preferred by sign language users for visual quality and comprehensibility in more than 80% of comparisons.

View source

Similar papers

Sep 2026

Automated sign language digitisation: Leveraging computer vision and archived broadcast footage for accessible avatar generation

Sign language serves as the primary linguistic modality for the deaf and hard-of-hearing community. However, the vast majority of sign language content contained within historical video archives remains inaccessible — unindexed, unsearchable and effectively invisible to computational analysis. This paper proposes a nov...

Takashi Koyano · 0 citations
Aug 2026

Variational Sign Language Translation

A novel framework based on conditional Variational autoencoder for SLT (VSLT) that facilitates direct and sufficient cross-modal alignment between sign language videos and spoken language text is proposed, and a shared Attention Residual Gaussian Distribution (ARGD) which considers the textual information as a residual...

Rui Zhao, Liang Zhang, Biao Fu et al. · 0 citations
#computer vision Preprint Sep 2026

SignFLIP: A Unified Model for Sign Language Translation and Generation via Stage-wise Alignment at Scale

Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unifi...

Zhaoyi An, Si-Han Tan, Youngbae Hwang et al. · 0 citations
Preprint Sep 2026

SignSeek: Learning Transferable Representations for Sign Dictionary Retrieval

SignSeek sets a new state-of-the-art performance in cross-corpus retrieval on ASL-Citizen, WLASL, and NMFs-CSL without any downstream fine-tuning, surpassing methods explicitly trained on BSL and outperforming prior skeleton-based methods.

Sobhan Asasi, Ozge Mercanoglu Sincan, Richard Bowden · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.