Preprint
Aug 2026
Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer
This work shows that a base T2AV model can be turned into a voice-cloning model by adding a single zero-initialized linear layer on top of its audio backbone, fine-tuning for a comparatively short training schedule, and conditioning on a short reference recording at inference time.
Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko et al.
· 0 citations