SAGE: Semantic Audio Generative Encoder
This paper introduces SAGE, Semantic Audio Generative Encoder: a compact variational autoencoder that shapes its latent by distilling embeddings from a pretrained audio-text model, combining high reconstruction fidelity, state-of-the-art semantic structure, and fast inference.
Francesco Brigante, Luca Cerovaz, Davide Marincione et al.
· 0 citations