BreezyVoice adapts CosyVoice to Taiwanese Mandarin with g2pW-based phonetic augmentation and a two-stage iconic-unit voice cloning pipeline, improving pronunciation accuracy and cloning robustness.
A Study on Incorporating Whisper for Robust Speech Assessment
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
This research introduces an enhanced version of the multi-objective speech assessment model--MOSA-Net+, by leveraging the acoustic features from Whisper, a large-scaled weakly supervised model. We first investigate the effectiveness of Whisper in deploying a more robust speech assessment model. After that, we explore combining representations from Whisper and SSL models. The experimental results reveal that Whisper's embedding features can contribute to more accurate prediction performance. Moreover, combining the embedding features from Whisper and SSL models only leads to marginal improvement. As compared to intrusive methods, MOSA-Net, and other SSL-based speech assessment models, MOSA-Net+ yields notable improvements in estimating subjective quality and intelligibility scores across all evaluation metrics in Taiwan Mandarin Hearing In Noise test - Quality & Intelligibility (TMHINT-QI) dataset. To further validate its robustness, MOSA-Net+ was tested in the noisy-and-enhanced track of the VoiceMOS Challenge 2023, where it obtained the top-ranked performance among nine systems.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
BreezyVoice: Adapting TTS for Taiwanese Mandarin with Enhanced Polyphone Disambiguation -- Challenges and Insights
BreezyVoice adapts CosyVoice to Taiwanese Mandarin with g2pW-based phonetic augmentation and a two-stage iconic-unit voice cloning pipeline, improving pronunciation accuracy and cloning robustness.