Adding white-noise copies of five minutes of a target speaker's data and rebalancing sampling lets a small ForwardTacotron model beat zero-shot baselines on speaker similarity with only four high-resource speakers.
FastSpeech: Fast, robust and controllable text-to-speech,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron
Adding white-noise copies of five minutes of a target speaker's data and rebalancing sampling lets a small ForwardTacotron model beat zero-shot baselines on speaker similarity with only four high-resource speakers.