A code-mixing guided preference-learning method for TTS produces synthetic data that lowers mixed error rate when fine-tuning Whisper on the SEAME Mandarin-English corpus.
Text-to-speech data augmentation for low resource speech recognition
3 Pith papers cite this work. Polarity classification is still indexing.
abstract
Nowadays, the main problem of deep learning techniques used in the development of automatic speech recognition (ASR) models is the lack of transcribed data. The goal of this research is to propose a new data augmentation method to improve ASR models for agglutinative and low-resource languages. This novel data augmentation method generates both synthetic text and synthetic audio. Some experiments were conducted using the corpus of the Quechua language, which is an agglutinative and low-resource language. In this study, a sequence-to-sequence (seq2seq) model was applied to generate synthetic text, in addition to generating synthetic speech using a text-to-speech (TTS) model for Quechua. The results show that the new data augmentation method works well to improve the ASR model for Quechua. In this research, an 8.73% improvement in the word-error-rate (WER) of the ASR model is obtained using a combination of synthetic text and synthetic speech.
representative citing papers
A Qwen3-ASR-based two-speaker, 21-language transcription system cuts its official error metric from 30.53 to 23.70 on the MLC-SLM 2026 dev set; supervised fine-tuning delivers most of the gain.
The authors propose and test a data augmentation framework based on deepfake audio to improve training of speech-to-text transcription models.
citing papers explorer
-
Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech
A code-mixing guided preference-learning method for TTS produces synthetic data that lowers mixed error rate when fine-tuning Whisper on the SEAME Mandarin-English corpus.
-
Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech
A Qwen3-ASR-based two-speaker, 21-language transcription system cuts its official error metric from 30.53 to 23.70 on the MLC-SLM 2026 dev set; supervised fine-tuning delivers most of the gain.
-
Deepfake audio as a data augmentation technique for training automatic speech to text transcription models
The authors propose and test a data augmentation framework based on deepfake audio to improve training of speech-to-text transcription models.