Skip to content

Author

Dhiraj Sunehra

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Embedding-Guided Neural Voice Conversion for Indian Regional-Language Speech Transformation

Voice conversion (VC) is an emerging technology in speech processing that aims to modify an utterance so that it sounds as though it were spoken by another speaker while preserving its linguistic content. High-quality voice conversion has broad applications, including speech synthesis, assistive communication, and entertainment applications such as multilingual dubbing. However, current embedding-guided voice conversion (EGVC) frameworks often struggle with generalization and naturalness under regional data-scarcity conditions. This study explores these limitations by evaluating an EGNVC framework adapted for low-resource regional-language pairs. The proposed framework incorporates the Harvest pitch-extraction algorithm alongside pretrained speaker representations to guide cross-gender pitch transitions while attempting to preserve speaker-identity profiles. Experimental results show that, although the framework successfully shifts macro-level pitch contours across genders, spectral alterations lead to substantial acoustic distortion and reduced intelligibility. Specifically, Kannada speech conversion achieved a localized objective intelligibility score of STOI = 0.12, whereas Malayalam transformations exhibited substantial spectral variation, with an MCD of 169.75, highlighting significant language-specific barriers to regional voice conversion. Kannada achieved higher intelligibility (STOI = 0.12) than Malayalam, whereas Malayalam required greater spectral modification (MCD = 169.75), indicating language-specific challenges in voice conversion. These baseline metrics delineate the empirical limitations of current embedding-guided architectures for Dravidian languages and indicate that substantial advances in spectral mapping are required before such systems can be integrated into real-time assistive or localized voice-synthesis applications.

B. A, Singh S. P., Dhiraj Sunehra · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.