A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies
This paper provides an evaluation and design framework for comparing representation-model pairs and shows that RVQ's residual order gives ordered capacity but not ordered semantics, and that AudioLM's semantic-versus-acoustic cascade is one explicit placement of this boundary rather than a universal template.