Author

Yifan Zhao

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Jul 2026

Generation in Generation: Fluid Co-Speech Gesture Synthesis With Generative Continuous Quantization

Motion quantization codebooks have been widely adopted to facilitate co-speech motion generation. However, the conventional quantization-based generation paradigm—which relies on probabilistic token sampling from limited discrete codebooks—suffers from two major limitations: crude, unreasonable motion representations and fixed, homogenized motion token sequences. To overcome these issues, we propose a novel explicit generation paradigm based on generative continuous quantization. Specifically, we first introduce a continuous quantization method to derive a set of generative motion units. This approach enables smoother and more accurate representation of human motion compared to classical methods. Building on these generative units, we further propose a compositional weight generation paradigm that replaces probabilistic sampling with deterministic, explicit motion synthesis. Moreover, as generalization capability is crucial for real-world deployment, we design a fully audio-aware encoder to extract style features that are decoupled from content. These features are integrated into the motion decoder via Adaptive Instance Normalization to enhance cross-speaker facial style generalization. Our method achieves state-of-the-art performance on two public datasets. Notably, owing to its concise and efficient architecture, our model attains an inference speed exceeding 4000 fps on the SHOW dataset, demonstrating strong potential for practical real-time applications.

Jialu Li, Yifan Zhao, Xin Guo et al. · 0 citations