Full-duplex voice assistants must continuously listen while speaking, handling user interruptions under low-latency and resource-constrained streaming conditions. Existing end-to-end full-duplex models can compromise reasoning-related capabilities after speech-domain adaptation, whereas cascaded pipelines introduce ext...
Zhi-Wei Lin, Tian-Jiao Du, Qiao-Chu Huang et al.· 0 citations
We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient acoustic conditions, voice characteristics, and o...
Lianghua Huang, Zhigang Wu, Yupeng Shi et al.· arXiv.org· 2 citations
X2Streaming-ASR is proposed, which decomposes streaming recognition into when to commit and what to commit, and achieves the best streaming CER among the evaluated systems on AISHELL-1 and AISHELL-3 with substantially lower latency.
Zhi-Wei Lin, Kaiqi Fu, Rime Wen et al.· 0 citations
Diff-Symbo is presented, an innovative method that uses latent diffusion model (LDM) to generate high-quality, diverse and long-duration symbolic music and improves the duration and the compositional consistency of music generation through an autoregressive approach.
Zhi-Wei Lin, Jun Chen, Bo-Shi Tang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.