SignSeek sets a new state-of-the-art performance in cross-corpus retrieval on ASL-Citizen, WLASL, and NMFs-CSL without any downstream fine-tuning, surpassing methods explicitly trained on BSL and outperforming prior skeleton-based methods.
Abstract
Sign language dictionaries are essential resources for sign language learners, yet automatically retrieving a sign from a dictionary, given only a query video, remains a challenging problem due to the natural variability between signers. Existing sign representation learning methods are built for closed-set recognition, producing embeddings that do not generalise to the open-set, signer-independent setting that retrieval demands. \textbf{SignSeek} closes this gap by contrastively learning sign representations with saliency-guided articulator masking. A contrastive objective aligns same-gloss signs across signers, while our Articulator Saliency-Guided Masking (ASGM) pinpoints the single most critical articulator per sign. This drives two complementary objectives, a masked contrastive alignment (MAC) loss that sees the sign through a single articulator and a masked prediction (MAP) loss that reconstructs it in latent space from the surrounding spatio-temporal context. Pretrained on 266K samples ($\sim$5,700 glosses) across multiple sign languages, \textbf{SignSeek} sets a new state-of-the-art performance in cross-corpus retrieval on ASL-Citizen, WLASL, and NMFs-CSL without any downstream fine-tuning. Strikingly, it achieves zero-shot generalisation to an entirely unseen British Sign Language (BSL), surpassing methods explicitly trained on BSL, and transfers seamlessly to isolated sign recognition and subtitle alignment, outperforming prior skeleton-based methods.
A prototype-structured sign embedding space is learned from continuous video annotated with signs, where each learnable prototype corresponds to a sign class, enabling the matching between dictionary exemplars and continuous sign instances.
Ryan Wong, Youngjoon Jang, Liliane Momeni et al.· 0 citations
The results show that loosely aligned broadcast data can provide effective weak supervision for learning sign representations that capture both lexical content and temporal structure, and raises top-5 temporal localization mean IoU from 0.235 to 0.465.
Oğuz Akif Tüfekcioğlu, Ezgi Ekin, Mustafa Kaan Çevik et al.· 0 citations
Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks. However, fine-tuning LLMs for Gloss-Free Sign Language Translation (GFSLT) remains a challenge. In this paper, we investigate how to effectively adapt LLMs to the GFSLT task. We show that there are two key issues that need to be...
Shi-Wei Gan, Xiao Liu, Ya-Feng Yin et al.· 1 citation
To the authors' knowledge this is the first description-based, open-vocabulary sign lookup from continuous signing without gloss supervision, and the first for JSL.
Santiago Poveda-Gutiérrez, Hideki Nakayama, Mayumi Bono· 0 citations
SignDino, a self-supervised sign-video encoder that moves the DINOv3 student--teacher recipe from the spatial domain of image crops to the temporal domain of tracked sign streams, provides a strong public self-supervised representation and shows competitive or state-of-the-art performance under matched downstream evalu...
Jun-Yi Hu, Zhe-Wen He, Hao Huang et al.· 0 citations
This work proposes a novel text-to-sign translation based on model pretraining, which enhances semantic alignment by inheriting codebook-oriented prior knowledge from masked self-supervised models.
Ninlawat Phuangchoke, C. Polprasert· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.