Progress in dialectal speech technology is hindered by the scarcity of large-scale, real-world corpora. For Minnan speech, existing resources remain limited, and few provide paired Minnan and Mandarin transcripts at scale. To address these gaps, we introduce WenetSpeech-Min, an open-source corpus comprising around 10,0...
Hao-Yu Zhang, Chun-Jiang He, Hong-Tao Li et al.· 0 citations
Recent advancements in discrete token-based speech generation have highlighted the importance of efficient token-to-waveform synthesis in streaming and dialogue scenarios. Flow-matching acoustic decoders achieve high-quality token-to-mel generation, but their iterative sampling requires multiple neural function evaluat...
Han-Ke Xie, Xia-Ming Ren, Qi-Rui Zhan et al.· 0 citations
SemBridge is proposed, a training-only semantic-token anchoring framework for continuous-latent autoregressive speech generation and demonstrates that explicit semantic-token supervision for autoregressive state learning is an effective and general direction for continuous speech generation.
Han-Ke Xie, Haopeng Lin, J. Qian et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.