SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting
This work proposes SMART, an MLLM-guided temporal alignment framework for joint sign recognition and spotting that incorporates CSFormer, a CSLR-guided spotting module that injects recognition-derived gloss evidence into a boundary-aware spotting network.
Eunjee Choi, J. Sung, Seongwhan Cho et al.
· 0 citations