A reinforcement-learning model with a continuous action space that integrates sound synthesis during training and live performance is introduced and five complementary reward functions for matching target sounds are proposed.
The proposed RL-based dynamic control system successfully transforms score elements into performance actions which enable robots to deliver expressive music performances.
Audio-video foundation models trained at scale implicitly encode a vast repertoire of perceptual and physical knowledge: motion, identity, environmental sound, lip dynamics, light, and material. The practical bottleneck is no longer what such a model can synthesize, but what a user can ask of it. This course presents m...
Naomi ken korem, Matan Ben Yosef· Proceedings of the Special I...· 0 citations
Vocalized audio synthesis, the task of generating audio in which intelligible speech is embedded within an environmental soundscape, underpins applications such as podcast production and video dubbing, and VoxAudio, a causal autoregressive flow matching model that addresses this problem from three complementary aspects...
Wenxiang Guo, Changhao Pan, Ziyue Jiang et al.· 0 citations
Recent audio generation systems have progressed from single-modality synthesis to generating complex acoustic scenes containing speech, music, and sound effects. Therefore, evaluating these models requires assessing multiple interacting capabilities, including semantic fidelity, speaker consistency, and temporal contro...
Zihao Zheng, Xuenan Xu, Jiahao Mei et al.· 0 citations
StrixAE, an agent based on a multimodal large language model (MLLM), outperforms most existing open-source and proprietary solutions, achieving state-of-the-art performance across multiple perceptual metrics and demonstrating strong generalization robustness.
Cheng-Lin Wu, Jun-Jie Wu, Jin-Hang Chen et al.· 0 citations
Synthesizer inversion is challenging for two main reasons: 1) Distinct parameter configurations can produce perceptually similar sounds. 2) Parameter-space losses often fail to reflect rendered audio similarity, while the synthesizer being a non-differentiable black box prevents simple audio-domain supervision. To addr...
Tristan Wu, Daniel Chin, Ju-Nan Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.