Skip to content

Author

Zhe-Wen He

We have 2 of 4 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#computer vision Preprint Sep 2026

SignDino: Self-Supervised Sign Language Representation Learning via Temporal-Axis Self-Distillation

Self-supervised sign language representation learning must model two properties not central to natural-image SSL: signs are produced by a small set of anatomically distinct articulators, and their meaning depends on the temporal organisation of those articulators. We introduce SignDino, a self-supervised sign-video encoder that moves the DINOv3 student--teacher recipe from the spatial domain of image crops to the temporal domain of tracked sign streams. Each video is decomposed into left-hand, right-hand, and face streams by a detector-first YOLOv8n+ByteTrack pipeline. A frozen DINOv3 ViT-B/16 embeds each per-frame anatomical crop, while lightweight temporal Transformers, not the image backbone, form the student and EMA teacher. They are trained by temporal DINO self-distillation, frame-level masked-token prediction in the style of iBOT, KoLeo feature spreading, and Gram anchoring of the frame-to-frame similarity structure. This design keeps strong image-level visual primitives fixed and learns only how articulator states evolve across time. We evaluate on sign-to-English translation, isolated sign recognition, and fingerspelling detection benchmarks. Across these tasks, SignDino provides a strong public self-supervised representation and shows competitive or state-of-the-art performance under matched downstream evaluation.

Jun-Yi Hu, Zhe-Wen He, Hao Huang et al. · 0 citations
Jul 2026

VTaMo: Video-Text Alignment Model for Sign Language Translation

VTaMo is presented, a framework that introduces explicit multi-granularity alignment at three levels: local alignment via entropy-regularized optimal transport with a learnable null token for fine-grained frame-to-token correspondences; global alignment via a learnable orthogonal transformation that calibrates embedding space geometry through Earth Mover's Distance.

Junyi Hu, Zhewen He, Hao Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.