Jul 2026
TokAN: Accent Normalization Using Self-Supervised Speech Tokens
TokAN is presented, a token-based accent normalization framework that operates on self-supervised discrete speech tokens extracted from a L1-L2 jointly trained vector-quantization (VQ) tokenizer, without the need of synthetic supervisory speech.
Qibing Bai, Shuai Wang, Yuhang Du et al.
· arXiv.org · 0 citations