Compressing Streaming Neural Audio Encoders via Latent-Space Distillation
This work studies how to compress a tokenizer by distillation, taking as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes - the last representation the two token interfaces share.
Prasanth Yadla, Mohammad Samragh, Dongseong Hwang et al.
· 0 citations