SignMimic achieves state-of-the-art-level performance on video quality, identity similarity, and frame continuity while also achieving minimal loss when performing back translation (SLT) on generated videos.
Abstract
We study the challenge of sign language video mimicking: given a driving video and a single reference frame, synthesize a video where the target signer reproduces the source motion while preserving identity and linguistic form. Prior pipelines entangle rigid motion, non-rigid deformation, and view-dependent completion in a monolithic generator, causing handshape drift and spatio-temporal instability. We present SignMimic, which (i) applies a TNet-based model to study SE(3) rigid canonicalization to stabilize global pose, (ii) performs non-rigid adaptation in a canonical space to preserve fine-grained articulators (hands/face) and coarticulation via NIF2D, and (iii) uses Pose-MAE-style completion before conditional video diffusion. This factorization injects geometric and linguistic priors, yielding shape and spatio-temporal consistency. On several large-scale datasets (ASL 50K, How2Sign, CSL News), SignMimic achieves state-of-the-art-level performance on video quality, identity similarity, and frame continuity while also achieving minimal loss when performing back translation (SLT) on generated videos. Ablations confirm the role of rigid canonicalization, non-rigid adaptation, and completion. Code is available at https://anonymous.4open.science/r/UniSignMimicTurbo-6088; model checkpoints and video examples will be released.
SignRefine is proposed, a sign language video generation model that produces comprehensible signing from 2D keypoint conditioning alone, generalizing across appearances and visual conditions, and is preferred by sign language users for visual quality and comprehensibility in more than 80% of comparisons.
A. Pelykh, Edward Fish, Ozge Mercanoglu Sincan et al.· 0 citations
Sign language production (SLP) aims to generate continuous signing motion from spoken language, often through gloss-to-pose generation. Prior work mainly follows two paradigms. Generative models synthesize motion from a learned prior or from noise, without reference to an observed signing instance, making rare hand con...
Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approaches rely on explicit motion representations such as 2D skelet...
Sangeyl Lee, Seunghyun Shin, S. Park et al.· 0 citations
Sign language-text alignment remains a fundamental challenge for text-driven sign language understanding. Existing methods predominantly rely on appearance-heavy RGB representations, which entangle motion semantics with visual variations and lead to ambiguous motion-language grounding. In this paper, we reformulate sig...
Chang Liu, Ke Han, Davide Talon et al.· 0 citations
Sign-language interpretation in public communication relies on qualified professional interpreters and can be difficult to scale, motivating robotic signing as a complementary accessibility interface. We present RoBoSTAR, a text-conditioned sign language production (SLP) framework for generating human-centric sign moti...
Yu-Jia Zeng, Chensheng Peng, Yu-Xin Chen et al.· 0 citations
This paper addresses static text-to-pose synthesis. Given a natural language description of a body configuration, the task is to generate the corresponding single-frame 3D human pose as a set of J root-relative joint positions in ℝJ×3. Discrete latent representations have been applied successfully to text-to-motion gen...