Jul 2026· PHM Society European Conference· Vol 9, pp. 1-12· 0 citations· 30 references
TL;DR
Results show consistent gains from lightweight domain adaptation on both held-out synthetic data and real-world recordings, confirming that synthetic data generation combined with LoRA-based fine-tuning is an effective and computationally practical strategy for improving ASR accuracy in specialized technical domains where labeled speech is scarce.
Abstract
Automatic speech recognition (ASR), or speech-to-text (STT), is becoming an important interface for AI systems in diagnostic workflows, but general-purpose ASR models often degrade in specialized technical domains. In diagnostic applications such as fault identification, root cause analysis, and repair recommendation, general-purpose ASR systems struggle with domain-specific terminology, abbreviations, part identifiers, and measurement expressions, leading to elevated transcription errors. This work presents a domain adaptation pipeline that unifies three components: a synthetic benchmarking framework in which domain-specific technical text is converted to speech via text-to-speech~(TTS) synthesis and transcribed by open-source ASR models to establish baseline performance; Low-Rank Adaptation~(LoRA)-based fine-tuning of Whisper Large-v3 using those synthetic audio-text pairs; and transfer validation on curated real-world automotive YouTube recordings to assess generalization beyond synthetic conditions. Using automotive technical language as a representative diagnostic domain, a data-scaling study employing progressively larger subsets of in-domain training data evaluates performance on a held-out test set via word error rate~(WER), character error rate~(CER), normalized error metrics, alphanumeric error rate, semantic similarity, and Bidirectional Encoder Representations from Transformers Score~(BERTScore). Results show consistent gains from lightweight domain adaptation on both held-out synthetic data and real-world recordings, confirming that synthetic data generation combined with LoRA-based fine-tuning is an effective and computationally practical strategy for improving ASR accuracy in specialized technical domains where labeled speech is scarce.
End-to-end on-device Automatic Speech Recognition (ASR) systems have demonstrated remarkable accuracy and efficiency in recent years. However, challenges persist in correctly transcribing infrequent named entities (e.g., geographical locations, business entities, person names, etc.) and handling diverse user accents, which remain underrepresented in training datasets. While information retrieval augmentation or Retrieval Augmentation Generation (RAG) for correction of named entities has shown promise in knowledge-grounded NLP tasks when paired with large language models (LLMs), its integration into real-time on-device systems is non-trivial due to computational constraints. We introduce a novel lightweight method combining phonetic-aware retrieval, vector-based semantic search and generative correction. The system leverages a lightweight phonetic index for rapid candidate entity retrieval and a dense vectorembedding module to refine predictions as well as model the ASR error output distribution in generative space. Additionally, we introduce a novel approach to model ASR errors in natural language. Experiments on test sets emphasizing place names, monuments, airports, and landscapes yielded an increase in correct Named Entity (NE) recognition accuracy by 9.6% compared to baseline. These gains underscore the efficacy of hybrid retrieval-generation paradigms in resource-constrained environments.
Kiranmayi Gandikota, Anunay Katare, C. Pandey et al.· International Conference on...· 0 citations
A unified phoneme-based TTS-to-ASR augmentation pipeline built around a multilingual TTS model trained from scratch using the F5-TTS architecture with language-ID conditioning is presented and phoneme-frequency-guided selection (PFGS) is proposed, which ranks candidate sentences using phoneme frequencies estimated from real ASR training labels.
Zhen Wang, Tian-Rui Wu, Rong-Qi Han et al.· 0 citations
Experimental results show that VALL-E outperforms the state-of-the-art zero-shot TTS system in terms of speech naturalness and speaker similarity and could preserve the speaker’s emotion and acoustic environment from the prompt in synthesis.
Speech Representation Fr'echet Distance loss (SR-FD), a training-time distributional regularizer for tokenizer-free flow-matching autoregressive TTS, is proposed, an intelligibility-improving distributional regularizer for few-step TTS.
Ho-Lam Chung, Kuan-Po Huang, Bo-Ru Lu et al.· 0 citations
ParaASR is introduced, an ASR system that leverages Multi-Token Prediction (MTP) to let a 4B LLM decoder emit multiple tokens per forward step and shows that decoder scaling, low-latency inference, and long-context transcription need not be competing goals when future-token proposals are anchored by the acoustic signal and guarded by autoregressive verification.
Automatic speech recognition (ASR) systems have achieved high accuracy with transformer-based models, enabling deployment in critical applications. However, they remain vulnerable to adversarial manipulation, particularly in black-box settings where attacks must preserve perceptual naturalness. This work introduces GATAS, a black-box testing approach that generates failure inducing inputs by operating in the phoneme-level latent space of a text- to-speech model. Instead of perturbing waveforms directly, the approach interpolates latent representations to induce transcription errors while remaining within the manifold of natural speech. The attack is formulated as a multi-objective optimization problem balancing semantic divergence and perceptual quality. Our empirical evaluation against both white-box and black-box baselines shows that GATAS achieves a 98% success rate while producing lower distortion and higher perceptual quality, as confirmed by human studies. Despite operating without gradient access, GATAS remains competitive against white-box methods, highlighting that representation and perceptual alignment are more critical than access to model internals. Overall, our results demonstrate that untargeted latent-space optimization enables the efficient generation of realistic and effective test cases for ASR systems.
Yanis Xabier Wilbrand Pena, Oliver Weissl, Andrea Stocco· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.