A diffusion transformer with a training-time timbre shifter improves zero-shot voice conversion similarity and intelligibility, and extends to singing conversion with F0 conditioning.
Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Zero-shot Voice Conversion with Diffusion Transformers
A diffusion transformer with a training-time timbre shifter improves zero-shot voice conversion similarity and intelligibility, and extends to singing conversion with F0 conditioning.