Full-duplex speech models are trained to converse with a person, but they are increasingly made to converse with each other, in self-play data generation, agent societies, and model-based evaluation. In that loop no human absorbs a timing error: each model's turn-taking is the other's input. We ask what timing the loop...
Li-Chen Zhu, Yueqian Lin, Yi-Heng Wang et al.· 0 citations
DynaCore is presented, a unified architecture for efficient LLM serving via system-architecture co-design that substantially reduces service-level latency over quantization and reconfigurable accelerators, and proposes disaggregated quantization, applying dual-side quantization to prefill and weight-only quantization t...
Cong Guo, Chi-Yue Wei, Bo-Wen Duan et al.· 0 citations
This study addresses challenges with Vortex, an architecture compatible with systolic-array-based accelerators with minimal hardware overhead, bridging the gap between extreme compression and efficient inference, and proposes codebook-wise contextual sparsity to align with VQ execution.
Haoxuan Shan, Cong Guo, Bo-Wen Duan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.