VTaMo: Video-Text Alignment Model for Sign Language Translation
VTaMo is presented, a framework that introduces explicit multi-granularity alignment at three levels: local alignment via entropy-regularized optimal transport with a learnable null token for fine-grained frame-to-token correspondences; global alignment via a learnable orthogonal transformation that calibrates embedding space geometry through Earth Mover's Distance.