Skip to content

Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning

Jul 2026 · arXiv.org · Vol abs/2607.22304 · 0 citations · 30 references
Computer Science

TL;DR

It is found that training on cloned data outperforms raw cross-lingual transfer for depression and anxiety detection on real Japanese speech, suggesting voice cloning is a promising direction for augmenting clinical speech data in low-resource languages.

Abstract

Synthetic data augmentation in speech is common practice for linguistic tasks like ASR, but has seen far less work for paralinguistic ones, especially clinical tasks where labelled data is expensive and some patient groups are underrepresented. Voice cloning is one such augmentation approach, but is typically evaluated on speech intelligibility (WER) or speaker similarity (SS) rather than on downstream performance, and it remains unclear whether these preserve the paralinguistic signal such tasks depend on. We benchmark eight voice cloning models on five paralinguistic tasks across public and clinical datasets, showing most preserve signal with modest degradation. We then clone English clinical speech into Japanese and find that training on cloned data outperforms raw cross-lingual transfer for depression and anxiety detection on real Japanese speech, suggesting voice cloning is a promising direction for augmenting clinical speech data in low-resource languages.

View source

Similar papers

#natural language process... Preprint Aug 2026

Scaling phoneme-based TTS augmentation for ASR: A unified pipeline and controlled study

A unified phoneme-based TTS-to-ASR augmentation pipeline built around a multilingual TTS model trained from scratch using the F5-TTS architecture with language-ID conditioning is presented and phoneme-frequency-guided selection (PFGS) is proposed, which ranks candidate sentences using phoneme frequencies estimated from real ASR training labels.

Zhen Wang, Tian-Rui Wu, Rong-Qi Han et al. · 0 citations
Open access Sep 2026

Toolkit for acoustic–phonetic analysis of naturalistic speech data

A major limitation in the speech sciences is access to naturalistic data in experimental settings and the difficulty of translating laboratory designs to real-world contexts. Researchers studying speech production or perception often rely on in-lab recordings, which constrain the sociolinguistic contexts examined and limit ecological validity. The Toolkit for Acoustic–Phonetic Analysis (TAPA) is an open-source pipeline that automates the acquisition, transcription, speaker diarization, forced alignment, and per-segment acoustic analysis of naturalistic, single/multi-speaker audio. The current release supports vowel formant extraction, stop voice onset time, and fricative spectral moments. We demonstrate TAPA on the 2016 U.S. presidential debate, extracting nearly 33,000 segments from a 90-min recording, and validate each measurement type against hand-coded annotation. Vowel formant agreement with expert measurements was high (F1 r = 0.89, F2 r = 0.87). Stop VOT showed reliable aggregate means but poor per-token agreement (r =  − 0.04) because of a training–deployment mismatch in the neural VOT classifier. Fricative spectral standard deviation agreed strongly with hand-coded values overall (r = 0.80), and center of gravity agreed strongly for sibilants (/s/ r = 0.87, /ʃ/ r = 0.94), while non-sibilant moments diverged systematically. These findings suggest that TAPA can be used to increase access to naturalistic speech data and speed up the processing timeline with experts’ supervision.

Ethan Kutlu, Emerson Peters, Ciara Tapanes et al. · 0 citations
Preprint Aug 2026

Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts

While text-based hallucination detection has been extensively studied, spoken hallucination detection remains largely unexplored, particularly for low-resource languages. We present the first multilingual spoken hallucination benchmark comprising 12,013 news samples across English, Russian, and Kazakh with controlled hallucinations of three types and three severity levels. Samples comprise original articles and aligned hallucinated counterparts in text and audio. We complement the synthetic corpus with 290 fact-checked fake news items collected natively in Russian (225) and Kazakh (65), translated into the other language and rendered through the same TTS-ASR pipeline. We assess fine-tuned multilingual encoders and, in zero-shot in-context settings, multimodal decoder models on transcript-based versus direct audio processing. Transcript-based detection generally outperforms direct audio processing, with binary-task degradation for strong encoders tracking per-language ASR error. On real-world fakes, synthetic-trained detectors transfer strongly (macro-F1 0.82-0.88 on original text), while Russian provenance analysis reveals both veracity-related and model-dependent machine-style signals, quantifying a key confound in synthetic hallucination benchmarks.

Meruyert Aristombayeva, Jason Samuel Lucas, Chaewan Chun et al. · 1 citation
Open access Jul 2026

Embedding-Guided Neural Voice Conversion for Indian Regional-Language Speech Transformation

Voice conversion (VC) is an emerging technology in speech processing that aims to modify an utterance so that it sounds as though it were spoken by another speaker while preserving its linguistic content. High-quality voice conversion has broad applications, including speech synthesis, assistive communication, and entertainment applications such as multilingual dubbing. However, current embedding-guided voice conversion (EGVC) frameworks often struggle with generalization and naturalness under regional data-scarcity conditions. This study explores these limitations by evaluating an EGNVC framework adapted for low-resource regional-language pairs. The proposed framework incorporates the Harvest pitch-extraction algorithm alongside pretrained speaker representations to guide cross-gender pitch transitions while attempting to preserve speaker-identity profiles. Experimental results show that, although the framework successfully shifts macro-level pitch contours across genders, spectral alterations lead to substantial acoustic distortion and reduced intelligibility. Specifically, Kannada speech conversion achieved a localized objective intelligibility score of STOI = 0.12, whereas Malayalam transformations exhibited substantial spectral variation, with an MCD of 169.75, highlighting significant language-specific barriers to regional voice conversion. Kannada achieved higher intelligibility (STOI = 0.12) than Malayalam, whereas Malayalam required greater spectral modification (MCD = 169.75), indicating language-specific challenges in voice conversion. These baseline metrics delineate the empirical limitations of current embedding-guided architectures for Dravidian languages and indicate that substantial advances in spectral mapping are required before such systems can be integrated into real-time assistive or localized voice-synthesis applications.

B. A, Singh S. P., Dhiraj Sunehra · 0 citations
Preprint Sep 2026

Language Orthogonalization of Self-Supervised Speech Representations for Cross-lingual Parkinson's Detection

Self-supervised speech models (S3Ms) provide powerful representations for Parkinson's disease (PD) detection, making cross-lingual transfer attractive for languages lacking labeled patient speech. However, these representations also encode language identity, which can confound this transfer: without target-language PD speech, classifiers may separate languages rather than pathology, yielding high specificity but low sensitivity on target patients. We propose \emph{language orthogonalization}, a closed-form ridge residualization of S3M features against external VoxLingua107 language embeddings, fitted using only healthy-control (HC) speech. By removing language-predictable components while retaining pathology-related variation, it produces a less language-dependent geometry in which HC representations concentrate while PD representations disperse. Across five S3M backbones, three speech tasks, and three target languages, our method consistently improves cross-lingual PD-detection performance while correcting the high-specificity/low-sensitivity failure.

Minu Kim, Eunjung Yeo, Kwanghee Choi et al. · 0 citations
#machine learning Preprint Sep 2026

Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS

Modern TTS systems approach human quality for high-resource languages but degrade when clean speech data is scarce. Modern Greek exemplifies this, lacking the curated corpora behind state-of-the-art synthesis. We propose a data curation recipe that transforms audiobook recordings into TTS-ready data via WhisperX alignment and filtering. Then we fine-tune Parler-TTS (880M), a prompt-based multilingual model whose pre-training encodes phonetic priors transferable to Greek. During development, we find that LLM-generated style prompts introduce speaker drift at inference. Replacing them with deterministic prompts resolves this, and a speaker-specific LoRA stage trained on 3.5 h of single-speaker data anchors identity while updating ~5% of parameters. Our system achieves WER 10.7% (2.9 above the ASR floor), MOS-I 4.00 (vs. 4.36 human speech), and near-human speaker consistency (MOS-C 4.24 vs. 4.30), showing that robust single-speaker Greek TTS is achievable with limited curated data.

Georgios Syllas, Efthymios Georgiou, Kosmas Kritsis et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.