FlexSpeech is a zero-shot TTS system that predicts phoneme durations autoregressively, renders speech with flow matching, and applies direct preference optimization to durations for fast style transfer.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
FlexSpeech is a zero-shot TTS system that predicts phoneme durations autoregressively, renders speech with flow matching, and applies direct preference optimization to durations for fast style transfer.