A jointly trained speech tokenizer and flow-matching decoder achieve strong zero-shot TTS and voice conversion on English and Mandarin, with WER below ground truth in several tests.
Wavlm: Large-scale self-supervised pre- training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization
A jointly trained speech tokenizer and flow-matching decoder achieve strong zero-shot TTS and voice conversion on English and Mandarin, with WER below ground truth in several tests.