A fine-tuned TTS with Whisper-based filtering generates synthetic Quebec French speech that trains ASR models well when combined with 10-60 hours of real audio, though a 13-14% WER gap to real-data training remains.
We proposed methods to improve synthetic speech quality and demonstrated that optimizing the generation pro- cesses enhanced synthetic data quality
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Towards Improved Speech Recognition through Optimized Synthetic Data Generation
A fine-tuned TTS with Whisper-based filtering generates synthetic Quebec French speech that trains ASR models well when combined with 10-60 hours of real audio, though a 13-14% WER gap to real-data training remains.