Skip to content
Review Open access

Voice Cloning: A Survey of Zero-Shot and Controllable Speech Synthesis

2026 · IEEE Access · Vol 14, pp. 122879-122896 · 0 citations · 91 references
Computer Science

TL;DR

This survey examines zero-shot voice cloning through the linked views of representation, generation, control, and deployment, and identifies open problems in benchmark standardization, speaker-similarity assessment, multilingual low-resource performance, preference-aligned synthesis, and secure real-world use.

Abstract

This survey examines zero-shot voice cloning through the linked views of representation, generation, control, and deployment. Rather than treating recent systems as isolated milestones, we organize the literature around the technical decisions that shape modern speaker-conditioned synthesis: neural codec design, acoustic-token modeling, autoregressive and non-autoregressive decoding, diffusion and flow-matching objectives, multilingual conditioning, streaming constraints, preference alignment, controllability, and safety. We review representative systems including YourTTS, VALL-E, Voicebox, F5-TTS, MaskGCT, Seed-TTS, CosyVoice 2, MiniMax-Speech, GLM-TTS, Spark-TTS, and Qwen3-TTS, and we compare how these systems trade off naturalness, speaker similarity, latency, controllability, and deployment risk. The survey is unified by a single frame, which is a deployment-oriented reading of the codec language model era. That frame is delivered through four contributions. We provide a four-axis taxonomy across representation, generation, control, and deployment. We provide a qualitative architectural comparison of representative systems along axes that headline metrics obscure. We provide a structured treatment of evaluation reliability that maps automatic and human metrics to the quality dimensions they actually capture. We provide a consolidated safety and governance perspective that couples technical mechanisms with dataset licensing and a pre-deployment checklist. Three field-level shifts recur across these views. The first is the move from waveform or spectrogram prediction toward discrete acoustic-token generation. The second is the transition from offline quality-first models toward streaming and conversational architectures. The third is the growing need for trustworthy evaluation, watermarking, consent-aware deployment, and demographic robustness. We conclude by identifying open problems in benchmark standardization, speaker-similarity assessment, multilingual low-resource performance, preference-aligned synthesis, and secure real-world use.

Read PDF

Similar papers

Preprint Aug 2026

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents

CuteTTS is presented, a compact continuous-autoregressive TTS system that reconciles high-fidelity generation with the latency demands of real-time interaction and introduces guidance-step distillation, which absorbs classifier-free guidance and multiple solver steps into a single interval-conditioned student.

Yuqian Zhang, Yao Shi, Kexin Huang et al. · 0 citations
Preprint Jul 2026

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm

In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5~Hz low-frame-rate speech tokenizer for r...

Bajian Xiang, Cheng Wen, Han Zhao et al. · 5 citations · ⚡1
Preprint Aug 2026

Luna-TTS Family Technical Report

Luna-TTS Family, diffusion-language-model-based TTS systems pretrained on 1 million hours of speech across Chinese, English, Japanese, and Korean are proposed, which achieves the best results on most objective, model-based, and human-rated metrics for NVV and emotion control.

Feng Yin, Shuai Shi, Junjie Zheng et al. · 0 citations

Models are Zero-Shot Text

Experimental results show that VALL-E outperforms the state-of-the-art zero-shot TTS system in terms of speech naturalness and speaker similarity and could preserve the speaker’s emotion and acoustic environment from the prompt in synthesis.

Unknown authors · 0 citations
Preprint Aug 2026

Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling

Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity but typically require high-quality speaker prompts and either strip away or entangle the acoustic environment with speaker characteristics, limiting their real-world applicability. We present an extended DAIEN-TTS, an en...

Ye-Xin Lu, Xin Wang, Yang Ai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programmable data source to turn a short voice reference (...

Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.