Skip to content

Simultaneous Translation between Sign Languages

Sep 2026 · 0 citations · 21 references
Computer Science

TL;DR

This work presents, to their knowledge, the first simultaneous sign-to-sign (S2S) translation system, with two wait-k regimes: test-time wait-k inference applied directly to a full-sentence model, and a trained wait-k model via stochastic multi-path supervision.

Abstract

Deaf and hard-of-hearing (DHH) signers cannot converse in real time across different sign languages today: existing sign-to-sign translation systems run offline, requiring the full source clip before any target sign is emitted. Live use cases - e.g. broadcast interpretation and two-way video calls - instead demand simultaneous output, while the source signer is still signing. We present, to our knowledge, the first simultaneous sign-to-sign (S2S) translation system, with two wait-k regimes: test-time wait-k inference applied directly to a full-sentence model, and a trained wait-k model via stochastic multi-path supervision. We further introduce ca-Stream-AL, a computation-aware latency metric for streaming output. Averaged across six S2S directions on both a smaller human-verified test set and a larger synthetic S2S corpus, our streaming system achieves a 38% ca-Stream-AL reduction while staying within a 9% DTW-PA-MPJPE increase and a 2.1 BLEU-4 drop compared to the full-sentence baseline. A word-order case study probes how the streaming model handles word order mismatch between different sign languages - a consequence of simultaneous translation.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks

This work introduces an alternative inspired by language-learning assessment, using an open-weight-LLM QA protocol that measures salient content preservation that aligns more closely with human rankings and is six to seven times more paraphrase-invariant than BLEU-4.

Oline Ranum, Edward Fish, Simon Hadfield et al. · 0 citations
#computer vision Preprint Oct 2026

ReSCUE: Re-translation with Sentence Commitment for Unsegmented Long-Form Simultaneous Sign Language Translation

Simultaneous Sign Language Translation (SLT) is critical for real-time communication, yet existing methods remain largely confined to sentence-level, offline settings that assume pre-segmented inputs. These assumptions hinder deployment in realistic scenarios involving continuous, unsegmented video streams. We present...

Si-Han Ren, Gao-Zheng Li, Yuan-Shang Quan et al. · 0 citations
Conference Open access Sep 2026

A Gloss-driven Indian Sign Language Production System Using Learned Pose Representations

A scalable and modular SLP framework based on Sign-Pose-VQ-VAE model, designed for low-resource settings, achieves state-of-the-art performance among keypoint-based methods on the PHOENIX14T benchmark, attaining a BLEU-4 score of 10.03 and surpassing the previous best method by 0.67 points.

Suvajit Patra, Arkadip Maitra, Swami Punyeshwarananda et al. · 0 citations
Review Oct 2026

Machine Translation for Sign Languages

Sign language machine translation has progressed substantially over the past decade, evolving from isolated sign recognition to end-to-end translation systems. Advances in pose estimation, transformer architectures, and large-scale dataset collection have driven progress, yet challenges remain. Datasets are limited com...

Ozge Mercanoglu Sincan, A. Pelykh, Edward Fish et al. · 0 citations
#computer vision Preprint Sep 2026

Zero-Shot Cross-Lingual Recognition of Sign Language Handshapes

This work presents the first zero-shot cross-lingual framework for handshape recognition, transferring from ASL to Catalan Sign Language (LSC), and leverages the decomposition of handshapes into five phonological features shared across both languages, to decode LSC handshapes from predicted features via a composite pho...

Marcel Granero-Moya, Carolina del Corral Farrarós, G. Haro et al. · 0 citations
Conference Aug 2026

Real-Time Multilingual Speech-to-Text AR Captioning Glasses for the Deaf and Hard-of-Hearing

For people who are deaf or hard-of-hearing (HoH), everyday conversations can be difficult to navigate, even with modern hearing aids. Augmented reality (AR) smart glasses offer a practical way to provide live captions directly in the user’s field of view. However, current models are often too expensive, rely exclusivel...

Mahesh Paul J, Soundharesh M, Vaissnave V et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Jun 3, 2026

MIT researchers teach AI models to interpret charts

The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.